Microsoft’s multi-agent AI system tops Anthropic’s Mythos on cybersecurity benchmark
88.45%. Microsoft's multi-agent system just dethroned Anthropic on cybersecurity benchmarks—here's why the shift to agent-swarms over single models matters.

Why it matters
Microsoft's MDASH system demonstrates a fundamental shift in AI capability—moving from single foundation models to orchestrated multi-agent architectures. This has direct implications for how enterprises will evaluate and deploy AI for high-stakes security tasks, and signals that benchmark leadership now requires systems thinking, not just model scale.
The key facts
6 to knowMicrosoft MDASH scored 88.45% on CyberGym benchmark
Surpassed Anthropic's Mythos and OpenAI single-model systems
Uses 100+ specialized AI agents across multiple models
Cybersecurity vulnerability-scanning focus
Multi-agent orchestration as competitive differentiator
Published May 14, 2026
Go to the source
GeekWiregeekwire.com
Publisher excerpt: Microsoft's new vulnerability-scanning system, codenamed MDASH, scored 88.45% on the CyberGym benchmark, surpassing single-model systems from Anthropic and OpenAI by using more than 100 specialized AI agents across multiple models.