The Agent RaceJuly 9, 2026via The Decoder
OpenAI finds roughly 30 percent of popular AI coding test is broken
Why it matters
OpenAI's retraction of SWE-Bench Pro endorsement undermines a key capability metric used across the industry to compare models. This calls into question the validity of prior benchmark claims and forces a recalibration of how leaders should evaluate coding model performance.
Key signals
- OpenAI found ~30% of SWE-Bench Pro tasks are broken/invalid
- OpenAI withdrawing prior endorsement of the benchmark
- SWE-Bench Pro is widely used industry standard for measuring AI programming capability
- Benchmark integrity directly impacts model comparison credibility
- Published July 9, 2026 — recent discovery affecting current leaderboards
The hook
30 percent. That's how much of SWE-Bench Pro — the industry's go-to AI coding benchmark — OpenAI just declared broken.
OpenAI reviewed SWE-Bench Pro, a widely used test for measuring AI models' programming skills, and found roughly 30 percent of its tasks are broken. The company is pulling its earlier endorsement of the benchmark.