FrontierThe story, in brief

OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 with its latest API and two additional settings

OpenAI's GPT-5.6 Sol claims ARC-AGI-3 victory—but only outside the official benchmark. The asterisk matters.

Paper-cut illustration of a coral software window opening into a three-dimensional drafting space.
New tools for building and creating with AI.AI illustration by KeyNews
The KeyNews take

Why it matters

Benchmark gaming and API-dependent performance raise questions about fair capability comparison. When test conditions matter more than the model itself, practitioners need to know what they're actually evaluating.

The key facts

10 to know
  1. GPT-5.6 Sol: 38.3% on ARC-AGI-3 with OpenAI API; 7.8% in official test environment

  2. Anthropic Opus 5 holds official ARC-AGI-3 record (implied benchmark lead)

  3. OpenAI cites 'latest API features' and 'two additional settings' as differentiators outside official test

  4. ARC Prize test environment flagged as potentially using outdated API

  5. Benchmark setup neutrality questioned; OpenAI-Anthropic lab-race drama continues

  6. GPT-5.6 Sol: 38.3% on ARC-AGI-3 (with OpenAI API features and two additional settings)

  7. GPT-5.6 Sol: 7.8% on ARC-AGI-3 (official test environment)

  8. Anthropic's Opus 5 previously held ARC-AGI-3 record

  9. ARC Prize test environment claims provider-neutrality but may use outdated API

  10. Benchmark setup dispute between OpenAI and ARC Prize over fairness

Go to the source

The Decoderthe-decoder.com

Publisher excerpt: OpenAI counters Anthropic's ARC-AGI-3 record: GPT-5.6 Sol scores 38.3 percent, but only with its own API features instead of the official test setup, where the model landed at 7.8 percent. ARC Prize claims its test environment is provider-neutral, but may have used an outdated API that skewed the…
Read original report
Back to today's editionMore frontier news

Keep reading

Related stories

More from Frontier