FrontierThe story, in brief

Why we no longer evaluate SWE-bench Verified

OpenAI just declared SWE-bench Verified broken. Here's what that means for your AI eval strategy.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

OpenAI's public rejection of SWE-bench Verified—a widely-used coding benchmark—signals that frontier model evaluation is fragmenting. If the gold standard benchmark is compromised, how do you trust comparative claims about coding capability?

The key facts

5 to know
  1. OpenAI officially discontinued SWE-bench Verified evaluation

  2. Identified training data leakage and test contamination in benchmark

  3. Recommends migration to SWE-bench Pro as alternative

  4. Signals broader concern: frontier coding benchmarks may be unreliable for measuring real progress

  5. Published by OpenAI directly, not third-party analysis

Go to the source

OpenAI Blogopenai.com

Publisher excerpt: SWE-bench Verified is increasingly contaminated and mismeasures frontier coding progress. Our analysis shows flawed tests and training leakage. We recommend SWE-bench Pro.
Read original report
Back to today's editionMore frontier news

Keep reading

Related stories

More from Frontier