FrontierSeptember 5, 2026via The Decoder
Artificial Analysis overhauls its Intelligence Index after GPT-6 Astra scoring drew skepticism
Why it matters
Benchmark integrity under pressure: as frontier models get closer, how we measure them matters more. A major evaluation framework revising its methodology signals real debate about whether existing benchmarks can differentiate at the frontier.
Key signals
- Artificial Analysis Intelligence Index version 4.2 released
- GPT-6 Astra scores four points higher on new benchmark
- Astra still trails Claude Fable 5.1 on revised index
- Overhaul prompted by skepticism over Astra's original scoring
- Benchmark revision reflects frontier-lab capability measurement debates
- Artificial Analysis Intelligence Index updated to version 4.2
- GPT-6 Astra gains 4 points after revision
- Astra still trails Claude Fable 5.1 on the revised index
- Revision appears to address skepticism about Astra's original scoring
- Published September 5, 2026
The hook
Artificial Analysis rewrites its benchmark after GPT-6 Astra scores draw fire — and Astra still trails Claude.
Artificial Analysis has released version 4.2 of its Intelligence Index, likely in response to criticism that its benchmarks failed to capture GPT-6 Astra's actual progress. Astra now scores four points above its predecessor but still trails Anthropic's Claude Fable 5.1.