Why SWE-bench Verified no longer measures frontier coding capabilities
OpenAI just declared SWE-bench Verified obsolete. Here's what that means for how you measure AI coding progress.

Why it matters
OpenAI's decision to stop using SWE-bench Verified signals a major shift in how frontier AI capabilities are evaluated—and suggests the benchmark no longer represents real-world coding challenges that matter to builders.
The key facts
5 to knowOpenAI officially discontinuing SWE-bench Verified as evaluation metric
Published April 26, 2026 on OpenAI's official blog
54 points on Hacker News with 45 comments (moderate community engagement)
Benchmark saturation/ceiling effect implied—frontier models likely maxing out on test
Signals shift in coding capability measurement standards across industry
Go to the source
Hacker Newsopenai.com
Publisher excerpt: Article URL: Comments URL: Points: 54 # Comments: 45