How We Broke Top AI Agent Benchmarks: And What Comes Next
Berkeley researchers just exposed how top AI agent benchmarks break under real-world conditions.

Why it matters
Academic research challenging the validity of widely-cited AI agent benchmarks raises questions about how the industry actually measures progress and what it means for model comparisons that investors and leaders rely on.
The key facts
5 to knowBerkeley RDI published benchmark vulnerability analysis
Research identifies failures in top AI agent benchmarks
169 HN points and 44 comments indicate strong community interest
Published April 11, 2026
Focuses on trustworthiness and reliability of benchmark methodologies
Go to the source
Hacker Newsrdi.berkeley.edu
Publisher excerpt: Article URL: Comments URL: Points: 169 # Comments: 44