WorkMay 10, 2026via The Decoder

METR says it can barely measure Claude Mythos, Palo Alto Networks warns of autonomous AI attackers

Why it matters

As AI models advance faster than evaluation methods, safety benchmarks are becoming dangerously blind spots—a critical governance gap that investors and CTOs need to understand.

Key signals

  • METR: Only 5 out of 228 tasks in current test suite cover Claude Mythos's relevant capability range (2.2% coverage)
  • Palo Alto Networks: Frontier models autonomously chain vulnerabilities from initial access to data exfiltration in 25 minutes
  • Evaluation methods growing slower than model capability progression
  • Safety assessment blindness becoming systemic risk across frontier labs

The hook

METR can only measure 2% of Claude Mythos's capabilities. Meanwhile, frontier models are chaining exploits in 25 minutes.

METR can barely measure Claude Mythos Preview with its current test suite. Only five out of 228 tasks cover the relevant capability range. Meanwhile, Palo Alto Networks reports that frontier models autonomously chain vulnerabilities, shrinking the time from initial access to data exfiltration to jus

The week's key stories, every Friday.

ONE BRIEFING · EVERY FRIDAY · FREE

Free. Unsubscribe anytime.

METR says it can barely measure Claude Mythos, Palo Alto Networks warns of autonomous AI attackers | KeyNews.AI