FrontierSeptember 4, 2026via The Decoder

Benchmarks disagree on GPT-6 Astra, but its human-beating efficiency on ARC-AGI-3 pulls Chollet’s AGI forecast forward

Why it matters

GPT-6 Astra's human-level efficiency on ARC-AGI-3 signals a capability inflection that the frontier labs are measuring differently — and it's reshaping expert timelines for AGI. Practitioners need to know which benchmark to trust; enthusiasts are watching the lab-race narrative shift.

Key signals

  • GPT-6 Astra scores 169 points on Epoch AI benchmark
  • Artificial Analysis rates Astra no better than predecessor, behind Claude Fable 5.1
  • ARC-AGI-3: Astra achieves human-beating efficiency for the first time
  • François Chollet observes progress running 'twice as fast' as expected
  • Chollet moving up AGI forecast based on ARC-AGI-3 result
  • Benchmark disagreement: Epoch vs. Artificial Analysis contradiction on relative capability

The hook

Conflicting benchmarks, one clear winner: GPT-6 Astra beats humans on ARC-AGI-3 for the first time, and Chollet just moved his AGI forecast up.

OpenAI's GPT-6 Astra is drawing contradictory benchmark verdicts. Epoch AI puts it out in front with 169 points, while Artificial Analysis rates it no better than its predecessor and behind Claude Fable 5.1. The biggest surprise comes from ARC-AGI-3, where Astra works more efficiently than the avera

The week's key stories, every Friday.

For practitioners and enthusiasts — free, in your inbox.

Free forever. No spam.