Zyphra Releases ZAYA1-8B: A Reasoning MoE Trained on AMD Hardware That Punches Far Above Its Weight Class
760M active parameters. That's all Zyphra's ZAYA1-8B needs to outperform models 10x larger on math and coding.

Why it matters
Zyphra demonstrates that reasoning efficiency—not scale—is the new competitive edge. A sub-1B MoE model trained on AMD hardware is now competing with frontier closed models, signaling a major shift in how efficiency-focused builders will benchmark and deploy.
The key facts
8 to knowZAYA1-8B: 760M active parameters in MoE architecture
Outperforms open-weight models many times its size on math and coding benchmarks
Surpasses Claude 4.5 Sonnet on HMMT'25 benchmark
Competitive with DeepSeek-V3.2
Novel Markovian RSA test-time compute method for reasoning
Trained end-to-end on AMD Instinct MI300 hardware
Released under Apache 2.0 (open-weight)
Sets standard for intelligence density in small language model weight class
Go to the source
MarkTechPostmarktechpost.com
Publisher excerpt: Zyphra releases ZAYA1-8B, a reasoning Mixture of Experts model with only 760M active parameters that outperforms open-weight models many times its size on math and coding benchmarks — closing in on DeepSeek-V3.2 and surpassing Claude 4.5 Sonnet on HMMT'25 with its novel Markovian RSA test-time…