ChipsAugust 25, 2026via The Decoder
Nvidia says its Groq 3 LPX is four times faster than Cerebras, but the math is more complicated
Why it matters
Inference chip performance claims are reshaping the compute buildout: throughput headlines mask scaling requirements and architecture tradeoffs that practitioners must understand when designing AI infrastructure.
Key signals
- Groq 3 LPX: 3,400 tokens/sec on Gemma 4 31B
- Groq 3 LPX requires 64+ accelerators to achieve headline throughput
- Cerebras achieves similar speeds with 1-2 accelerators
- MoE model scaling behavior remains unresolved between architectures
- Chip entering full production
The hook
Nvidia's Groq 3 LPX hits 3,400 tokens/sec — but needs 64 accelerators to do it. Cerebras does it with two.
Nvidia is moving its Groq 3 LPX inference chip into full production and reports 3,400 tokens per second on Gemma 4 31B, four times faster than Cerebras. But the numbers don't tell the whole story. Nvidia needs at least 64 accelerators to get there, while Cerebras needs only one or two, according to …