New benchmark shows Claude Mythos and GPT-5.5 can develop real browser exploits autonomously
Claude Mythos just lapped GPT-5.5 on autonomous exploit development. There's a catch: it costs 12x more.

Why it matters
A new Carnegie Mellon benchmark reveals a meaningful capability gap between Claude Mythos and GPT-5.5 on a security-critical task (autonomous browser exploit development), but cost-performance tradeoffs are reshaping which model leaders actually deploy in production.
The key facts
5 to knowClaude Mythos outperforms GPT-5.5 on autonomous V8 engine exploit development
12x cost differential between Mythos and GPT-5.5
Benchmark tests real-world vulnerability exploitation capability
Carnegie Mellon University research
Agents-as-capability benchmark (autonomous agent performance)
Go to the source
The Decoderthe-decoder.com
Publisher excerpt: Researchers at Carnegie Mellon University built a new benchmark that measures how far AI agents can go when exploiting real vulnerabilities in Google's V8 engine. Mythos leads GPT-5.5 by a wide margin but costs twelve times as much.