k.Frontier
FrontierOpenAI Blog
KeyRank 78Why we no longer evaluate SWE-bench Verified
OpenAI's public rejection of SWE-bench Verified—a widely-used coding benchmark—signals that frontier model evaluation is fragmenting. If the gold standard benchmark is compromised, how do you trust comparative claims about coding capability?
Why it ranks · · OpenAI officially discontinued SWE-bench Verified evaluation · 2026-02-23
Read full story