FrontierSeptember 4, 2026via The Decoder
Benchmarks disagree on GPT-6 Astra, but its human-beating efficiency on ARC-AGI-3 pulls Chollet’s AGI forecast forward
Why it matters
GPT-6 Astra's human-level efficiency on ARC-AGI-3 signals a capability inflection that the frontier labs are measuring differently — and it's reshaping expert timelines for AGI. Practitioners need to know which benchmark to trust; enthusiasts are watching the lab-race narrative shift.
Key signals
- GPT-6 Astra scores 169 points on Epoch AI benchmark
- Artificial Analysis rates Astra no better than predecessor, behind Claude Fable 5.1
- ARC-AGI-3: Astra achieves human-beating efficiency for the first time
- François Chollet observes progress running 'twice as fast' as expected
- Chollet moving up AGI forecast based on ARC-AGI-3 result
- Benchmark disagreement: Epoch vs. Artificial Analysis contradiction on relative capability
The hook
Conflicting benchmarks, one clear winner: GPT-6 Astra beats humans on ARC-AGI-3 for the first time, and Chollet just moved his AGI forecast up.
OpenAI's GPT-6 Astra is drawing contradictory benchmark verdicts. Epoch AI puts it out in front with 169 points, while Artificial Analysis rates it no better than its predecessor and behind Claude Fable 5.1. The biggest surprise comes from ARC-AGI-3, where Astra works more efficiently than the avera…