OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 with its latest API and two additional settings
OpenAI's GPT-5.6 Sol claims ARC-AGI-3 victory—but only outside the official benchmark. The asterisk matters.

Why it matters
Benchmark gaming and API-dependent performance raise questions about fair capability comparison. When test conditions matter more than the model itself, practitioners need to know what they're actually evaluating.
The key facts
10 to knowGPT-5.6 Sol: 38.3% on ARC-AGI-3 with OpenAI API; 7.8% in official test environment
Anthropic Opus 5 holds official ARC-AGI-3 record (implied benchmark lead)
OpenAI cites 'latest API features' and 'two additional settings' as differentiators outside official test
ARC Prize test environment flagged as potentially using outdated API
Benchmark setup neutrality questioned; OpenAI-Anthropic lab-race drama continues
GPT-5.6 Sol: 38.3% on ARC-AGI-3 (with OpenAI API features and two additional settings)
GPT-5.6 Sol: 7.8% on ARC-AGI-3 (official test environment)
Anthropic's Opus 5 previously held ARC-AGI-3 record
ARC Prize test environment claims provider-neutrality but may use outdated API
Benchmark setup dispute between OpenAI and ARC Prize over fairness
Go to the source
The Decoderthe-decoder.com
Publisher excerpt: OpenAI counters Anthropic's ARC-AGI-3 record: GPT-5.6 Sol scores 38.3 percent, but only with its own API features instead of the official test setup, where the model landed at 7.8 percent. ARC Prize claims its test environment is provider-neutral, but may have used an outdated API that skewed the…