Monday, February 23, 2026

Start of archive·July 7
keynews.ai

Key AI news stories in enterprise and tech.

For practitioners and enthusiasts

Edition · 2026-02-23

k.Frontier
FrontierOpenAI Blog
KeyRank 78

Why we no longer evaluate SWE-bench Verified

OpenAI's public rejection of SWE-bench Verified—a widely-used coding benchmark—signals that frontier model evaluation is fragmenting. If the gold standard benchmark is compromised, how do you trust comparative claims about coding capability?

Why it ranks · · OpenAI officially discontinued SWE-bench Verified evaluation · 2026-02-23

Read full story
The six pillarsRanked by KeyRank · 2026-02-23

Frontier

20% share

Models, benchmarks, the lab race. · 1 story

  1. 1Why we no longer evaluate SWE-bench Verified78

Agents

quiet today

Autonomous systems in the wild. · 0 stories

  1. Nothing cleared the bar today.

Money

quiet today

Funding, M&A, valuations, earnings. · 0 stories

  1. Nothing cleared the bar today.

Chips

quiet today

Silicon, racks, the compute buildout. · 0 stories

  1. Nothing cleared the bar today.

Work

20% share

Jobs, industries, people & policy. · 1 story

  1. 1Import AI 446: Nuclear LLMs; China's big AI benchmark; measurement and AI policy72

Get it by email.

Daily · Weekly · Monthly