Anthropic's new benchmark claims Claude can match human experts in bioinformatics
Claude matches human experts on bioinformatics benchmarks. But here's what Anthropic isn't telling you.

Why it matters
Anthropic is using domain-specific benchmarking (BioMysteryBench) to demonstrate Claude's capability parity with human experts in high-stakes fields like bioinformatics. This signals a shift toward vertical capability claims rather than horizontal benchmarks—and raises questions about benchmark design and real-world deployment readiness.
The key facts
5 to knowNew benchmark: BioMysteryBench for bioinformatics domain evaluation
Claim: Claude matches human expert performance on benchmark tasks
Article notes important caveats—benchmark results come with qualifications
Domain-specific benchmarking as competitive positioning strategy
Published April 30, 2026—timing suggests recent capability release
Go to the source
The Decoderthe-decoder.com
Publisher excerpt: With BioMysteryBench, Anthropic wants to show that Claude can solve real bioinformatics problems at an expert level. The results are promising, but come with important caveats.