FrontierThe story, in brief

OpenAI finds roughly 30 percent of popular AI coding test is broken

30 percent. That's how much of SWE-Bench Pro — the industry's go-to AI coding benchmark — OpenAI just declared broken.

Paper-cut illustration of a coral software window opening into a three-dimensional drafting space.
New tools for building and creating with AI.AI illustration by KeyNews
The KeyNews take

Why it matters

OpenAI's retraction of SWE-Bench Pro endorsement undermines a key capability metric used across the industry to compare models. This calls into question the validity of prior benchmark claims and forces a recalibration of how leaders should evaluate coding model performance.

The key facts

5 to know
  1. OpenAI found ~30% of SWE-Bench Pro tasks are broken/invalid

  2. OpenAI withdrawing prior endorsement of the benchmark

  3. SWE-Bench Pro is widely used industry standard for measuring AI programming capability

  4. Benchmark integrity directly impacts model comparison credibility

  5. Published July 9, 2026 — recent discovery affecting current leaderboards

Go to the source

The Decoderthe-decoder.com

Publisher excerpt: OpenAI reviewed SWE-Bench Pro, a widely used test for measuring AI models' programming skills, and found roughly 30 percent of its tasks are broken. The company is pulling its earlier endorsement of the benchmark.
Read original report
Back to today's editionMore frontier news

Keep reading

Related stories

More from Frontier