FrontierThe story, in brief

I set 10 honesty traps for Claude Opus 4.8 - and a legal test broke it

Claude Opus 4.8 fails on legal reasoning. Here's what that means for enterprise deployments.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

Independent testing reveals capability gaps in Anthropic's latest model under adversarial conditions, raising questions about reliability in high-stakes domains like legal and finance where accuracy is non-negotiable.

The key facts

5 to know
  1. Claude Opus 4.8 tested against 4.7 across four domains: coding, medical, finance, legal

  2. Legal domain test specifically broke the model

  3. Testing methodology included 10 'honesty traps'

  4. Results cross-checked with multiple AI systems

  5. Published by ZDNet (third-party independent evaluation)

Go to the source

ZDNet AIzdnet.com

Publisher excerpt: I tested Opus 4.8 against 4.7 using coding, medical, finance, and legal traps, then cross-checked the results with multiple AIs.
Read original report
Back to today's editionMore frontier news

Keep reading

Related stories

More from Frontier