Why we no longer evaluate SWE-bench Verified
OpenAI just declared SWE-bench Verified broken. Here's what that means for your AI eval strategy.

Why it matters
OpenAI's public rejection of SWE-bench Verified—a widely-used coding benchmark—signals that frontier model evaluation is fragmenting. If the gold standard benchmark is compromised, how do you trust comparative claims about coding capability?
The key facts
5 to knowOpenAI officially discontinued SWE-bench Verified evaluation
Identified training data leakage and test contamination in benchmark
Recommends migration to SWE-bench Pro as alternative
Signals broader concern: frontier coding benchmarks may be unreliable for measuring real progress
Published by OpenAI directly, not third-party analysis
Go to the source
OpenAI Blogopenai.com
Publisher excerpt: SWE-bench Verified is increasingly contaminated and mismeasures frontier coding progress. Our analysis shows flawed tests and training leakage. We recommend SWE-bench Pro.