FrontierSeptember 5, 2026via The Decoder

Artificial Analysis overhauls its Intelligence Index after GPT-6 Astra scoring drew skepticism

Why it matters

Benchmark integrity under pressure: as frontier models get closer, how we measure them matters more. A major evaluation framework revising its methodology signals real debate about whether existing benchmarks can differentiate at the frontier.

Key signals

  • Artificial Analysis Intelligence Index version 4.2 released
  • GPT-6 Astra scores four points higher on new benchmark
  • Astra still trails Claude Fable 5.1 on revised index
  • Overhaul prompted by skepticism over Astra's original scoring
  • Benchmark revision reflects frontier-lab capability measurement debates
  • Artificial Analysis Intelligence Index updated to version 4.2
  • GPT-6 Astra gains 4 points after revision
  • Astra still trails Claude Fable 5.1 on the revised index
  • Revision appears to address skepticism about Astra's original scoring
  • Published September 5, 2026

The hook

Artificial Analysis rewrites its benchmark after GPT-6 Astra scores draw fire — and Astra still trails Claude.

Artificial Analysis has released version 4.2 of its Intelligence Index, likely in response to criticism that its benchmarks failed to capture GPT-6 Astra's actual progress. Astra now scores four points above its predecessor but still trails Anthropic's Claude Fable 5.1.

The week's key stories, every Friday.

For practitioners and enthusiasts — free, in your inbox.

Free forever. No spam.