OpenAI finds roughly 30 percent of popular AI coding test is broken
30 percent. That's how much of SWE-Bench Pro — the industry's go-to AI coding benchmark — OpenAI just declared broken.

Why it matters
OpenAI's retraction of SWE-Bench Pro endorsement undermines a key capability metric used across the industry to compare models. This calls into question the validity of prior benchmark claims and forces a recalibration of how leaders should evaluate coding model performance.
The key facts
5 to knowOpenAI found ~30% of SWE-Bench Pro tasks are broken/invalid
OpenAI withdrawing prior endorsement of the benchmark
SWE-Bench Pro is widely used industry standard for measuring AI programming capability
Benchmark integrity directly impacts model comparison credibility
Published July 9, 2026 — recent discovery affecting current leaderboards
Go to the source
The Decoderthe-decoder.com
Publisher excerpt: OpenAI reviewed SWE-Bench Pro, a widely used test for measuring AI models' programming skills, and found roughly 30 percent of its tasks are broken. The company is pulling its earlier endorsement of the benchmark.