FrontierThe story, in brief

New benchmark shows Claude Mythos and GPT-5.5 can develop real browser exploits autonomously

Claude Mythos just lapped GPT-5.5 on autonomous exploit development. There's a catch: it costs 12x more.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

A new Carnegie Mellon benchmark reveals a meaningful capability gap between Claude Mythos and GPT-5.5 on a security-critical task (autonomous browser exploit development), but cost-performance tradeoffs are reshaping which model leaders actually deploy in production.

The key facts

5 to know
  1. Claude Mythos outperforms GPT-5.5 on autonomous V8 engine exploit development

  2. 12x cost differential between Mythos and GPT-5.5

  3. Benchmark tests real-world vulnerability exploitation capability

  4. Carnegie Mellon University research

  5. Agents-as-capability benchmark (autonomous agent performance)

Go to the source

The Decoderthe-decoder.com

Publisher excerpt: Researchers at Carnegie Mellon University built a new benchmark that measures how far AI agents can go when exploiting real vulnerabilities in Google's V8 engine. Mythos leads GPT-5.5 by a wide margin but costs twelve times as much.
Read original report
Back to today's editionMore frontier news

Keep reading

Related stories

More from Frontier