FrontierJuly 26, 2026via The Decoder

Anthropic's Opus 5 blows past Fable 5 and GPT-5.6 Sol on the benchmark designed to measure real intelligence

Why it matters

Anthropic's latest model demonstrates a substantial capability leap on a reasoning-focused benchmark, with independent reflection abilities that haven't been observed in competing models. This signals a shift in the model capability race toward deeper logical reasoning as a differentiator.

Key signals

  • Claude Opus 5 scores 30.2% on ARC-AGI-3 benchmark
  • Previous record (GPT-5.6 Sol): 7.8%
  • 4x improvement over prior state-of-the-art
  • Model demonstrated independent reflection equation formulation—novel behavior
  • Benchmark designed to measure 'real intelligence' (reasoning-focused)
  • Competing models: Fable 5, GPT-5.6 Sol

The hook

30.2%. That's Anthropic's Claude Opus 5 on ARC-AGI-3—nearly quadrupling OpenAI's previous benchmark record.

Anthropic's Claude Opus 5 scored 30.2 percent on ARC-AGI-3, nearly quadrupling GPT-5.6 Sol's previous record of 7.8 percent. The benchmark's developers say the model independently formulated reflection equations, a behavior they had never seen from another model, and attribute to stronger logical reasoning.

The week's key stories, every Friday.

For practitioners and enthusiasts — free, in your inbox.

Free forever. No spam.