WorkThe story, in brief

UK's AI Security Institute finds standard benchmarks systematically underestimate what AI agents can actually do

Your AI benchmarks are lying to you. The UK's AI Security Institute just proved standard evals underestimate agent capabilities by 60%.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

Industry benchmark methodology is fundamentally flawed—token budget constraints mask true frontier progress. This has immediate implications for how founders, investors, and CTOs assess model capabilities and competitive positioning.

The key facts

5 to know
  1. UK AI Security Institute study covers 7 benchmarks

  2. Software engineering task success rates jumped ~25% with 10x token budget increase

  3. Frontier progress measured at 60% steeper than previous benchmarks suggested

  4. Newer models show largest gains from increased compute budget

  5. Finding: standard evaluations systematically underestimate agent capabilities

Go to the source

The Decoderthe-decoder.com

Publisher excerpt: In a study covering seven benchmarks, the UK's AI Security Institute shows that standard AI evaluations systematically underestimate agent capabilities by capping the compute budget. On software engineering tasks, success rates jumped about 25 percent when the token budget was increased tenfold.…
Read original report
Back to today's editionMore work news

Keep reading

Related stories

More from Work