I set 10 honesty traps for Claude Opus 4.8 - and a legal test broke it
Claude Opus 4.8 fails on legal reasoning. Here's what that means for enterprise deployments.

Why it matters
Independent testing reveals capability gaps in Anthropic's latest model under adversarial conditions, raising questions about reliability in high-stakes domains like legal and finance where accuracy is non-negotiable.
The key facts
5 to knowClaude Opus 4.8 tested against 4.7 across four domains: coding, medical, finance, legal
Legal domain test specifically broke the model
Testing methodology included 10 'honesty traps'
Results cross-checked with multiple AI systems
Published by ZDNet (third-party independent evaluation)
Go to the source
ZDNet AIzdnet.com
Publisher excerpt: I tested Opus 4.8 against 4.7 using coding, medical, finance, and legal traps, then cross-checked the results with multiple AIs.