FrontierJuly 26, 2026via The Decoder
Anthropic's Opus 5 blows past Fable 5 and GPT-5.6 Sol on the benchmark designed to measure real intelligence
Why it matters
Anthropic's latest model demonstrates a substantial capability leap on a reasoning-focused benchmark, with independent reflection abilities that haven't been observed in competing models. This signals a shift in the model capability race toward deeper logical reasoning as a differentiator.
Key signals
- Claude Opus 5 scores 30.2% on ARC-AGI-3 benchmark
- Previous record (GPT-5.6 Sol): 7.8%
- 4x improvement over prior state-of-the-art
- Model demonstrated independent reflection equation formulation—novel behavior
- Benchmark designed to measure 'real intelligence' (reasoning-focused)
- Competing models: Fable 5, GPT-5.6 Sol
The hook
30.2%. That's Anthropic's Claude Opus 5 on ARC-AGI-3—nearly quadrupling OpenAI's previous benchmark record.
Anthropic's Claude Opus 5 scored 30.2 percent on ARC-AGI-3, nearly quadrupling GPT-5.6 Sol's previous record of 7.8 percent. The benchmark's developers say the model independently formulated reflection equations, a behavior they had never seen from another model, and attribute to stronger logical reasoning.