ITBench-AA: Frontier Models Score Below 50% on the First Benchmark for Agentic Enterprise IT Tasks — by Artificial Analysis and IBM
Frontier models failing at 50% on enterprise IT tasks. The AI agents everyone's betting on can't handle your infrastructure yet.

Why it matters
New benchmark reveals a critical gap between enterprise AI capability claims and real-world IT operations performance. Frontier models score below 50% on agentic enterprise IT tasks, signaling that AI agent deployment in critical infrastructure remains immature despite industry hype.
The key facts
6 to knowITBench-AA is the first benchmark specifically designed for agentic enterprise IT tasks
Frontier models score below 50% on the benchmark
Co-developed by Artificial Analysis and IBM Research
Indicates significant gap between claimed AI capabilities and practical enterprise IT applications
Published May 27, 2026
Available on Hugging Face
Go to the source
Hugging Face Bloghuggingface.co

