The Agent RaceJuly 19, 2026via The Decoder
Moonshot's Kimi K3 outperforms Fable 5 in frontend code but lags far behind in complex math
Why it matters
Chinese AI models are now competitive on specialized tasks like frontend code generation, but fundamental capability gaps in mathematical reasoning remain. This signals a bifurcation in the global AI race: regional dominance on narrow benchmarks vs. broad reasoning leadership.
Key signals
- Moonshot Kimi K3 ranks #1 on Code Arena: Frontend (beats Claude Fable 5 and GPT-5.6 Sol)
- Kimi K3 scores ~39% on FrontierMath Tier 4 advanced math benchmark
- OpenAI and Anthropic models score ~90% on FrontierMath Tier 4
- First Chinese model to top Code Arena: Frontend rankings
- Math capability gap: 51 percentage point spread between Kimi K3 and top Western models
The hook
Moonshot's Kimi K3 just topped the Code Arena: Frontend rankings—beating Claude and GPT-5.6. But the math gap tells a different story.
Moonshot's Kimi K3 is the first Chinese model to top the Code Arena: Frontend rankings, beating Claude Fable 5 and GPT-5.6 Sol by a wide margin. But on advanced math, the gap is stark: Kimi K3 scores only about 39 percent on FrontierMath Tier 4, while models from OpenAI and Anthropic hit close to 90.