The Agent RaceJuly 9, 2026via The Decoder

OpenAI finds roughly 30 percent of popular AI coding test is broken

Why it matters

OpenAI's retraction of SWE-Bench Pro endorsement undermines a key capability metric used across the industry to compare models. This calls into question the validity of prior benchmark claims and forces a recalibration of how leaders should evaluate coding model performance.

Key signals

  • OpenAI found ~30% of SWE-Bench Pro tasks are broken/invalid
  • OpenAI withdrawing prior endorsement of the benchmark
  • SWE-Bench Pro is widely used industry standard for measuring AI programming capability
  • Benchmark integrity directly impacts model comparison credibility
  • Published July 9, 2026 — recent discovery affecting current leaderboards

The hook

30 percent. That's how much of SWE-Bench Pro — the industry's go-to AI coding benchmark — OpenAI just declared broken.

OpenAI reviewed SWE-Bench Pro, a widely used test for measuring AI models' programming skills, and found roughly 30 percent of its tasks are broken. The company is pulling its earlier endorsement of the benchmark.

The week's key stories, every Friday.

For practitioners and enthusiasts — free, in your inbox.

Free forever. No spam.

OpenAI finds roughly 30 percent of popular AI coding test is broken | KeyNews.AI