Even the latest AI models make three systematic reasoning errors, ARC-AGI-3 analysis shows
GPT-5.5 and Opus 4.7 both fail below 1% on ARC-AGI-3. Three systematic reasoning patterns explain why.

Why it matters
Latest frontier models exhibit consistent, identifiable reasoning failures on abstract reasoning tasks—revealing a capability ceiling that benchmarking data alone won't expose. This matters for teams betting on reasoning-as-AGI-path.
The key facts
5 to know160 game runs analyzed by ARC Prize Foundation
GPT-5.5 and Opus 4.7 both below 1% on ARC-AGI-3
Three systematic error patterns identified
Tasks solvable by humans with minimal effort
ARC-AGI-3 benchmark used as evaluation framework
Go to the source
The Decoderthe-decoder.com
Publisher excerpt: The ARC Prize Foundation analyzed 160 game runs of OpenAI's GPT-5.5 and Anthropic's Opus 4.7 on the ARC-AGI-3 benchmark. Three systematic error patterns explain why both models stay below 1 percent on tasks that humans can solve without much trouble.