FrontierThe story, in brief

Zyphra Releases ZAYA1-8B: A Reasoning MoE Trained on AMD Hardware That Punches Far Above Its Weight Class

760M active parameters. That's all Zyphra's ZAYA1-8B needs to outperform models 10x larger on math and coding.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

Zyphra demonstrates that reasoning efficiency—not scale—is the new competitive edge. A sub-1B MoE model trained on AMD hardware is now competing with frontier closed models, signaling a major shift in how efficiency-focused builders will benchmark and deploy.

The key facts

8 to know
  1. ZAYA1-8B: 760M active parameters in MoE architecture

  2. Outperforms open-weight models many times its size on math and coding benchmarks

  3. Surpasses Claude 4.5 Sonnet on HMMT'25 benchmark

  4. Competitive with DeepSeek-V3.2

  5. Novel Markovian RSA test-time compute method for reasoning

  6. Trained end-to-end on AMD Instinct MI300 hardware

  7. Released under Apache 2.0 (open-weight)

  8. Sets standard for intelligence density in small language model weight class

Go to the source

MarkTechPostmarktechpost.com

Publisher excerpt: Zyphra releases ZAYA1-8B, a reasoning Mixture of Experts model with only 760M active parameters that outperforms open-weight models many times its size on math and coding benchmarks — closing in on DeepSeek-V3.2 and surpassing Claude 4.5 Sonnet on HMMT'25 with its novel Markovian RSA test-time…
Read original report
Back to today's editionMore frontier news

Keep reading

Related stories

More from Frontier