WorkThe story, in brief

Agent skills look great in benchmarks but fall apart under realistic conditions, researchers find

34,000 skills tested. Almost none of them work in the real world. Your agent strategy needs a rethink.

Illustration of independent geometric mechanisms passing paper tasks along branching amber tracks.
AI agents and the coordination of work.AI illustration by KeyNews
The KeyNews take

Why it matters

New research reveals a critical gap between benchmark performance and real-world deployment of AI agent skills—a finding that directly challenges current assumptions about modular knowledge architectures and could reshape how companies approach agent capability stacking.

The key facts

4 to know
  1. Study tested 34,000 real-world AI agent skills

  2. Skills show strong benchmark performance but fail under realistic conditions

  3. Weaker models perform worse with skills than without them

  4. Gap between controlled evaluation and production deployment identified

Go to the source

The Decoderthe-decoder.com

Publisher excerpt: AI agents are supposed to tap into specialized knowledge through so-called skills, modular instructions they can pull up on the fly. But a study testing 34,000 real-world skills finds these enhancements barely help under realistic conditions. Weaker models actually perform worse with them than…
Read original report
Back to today's editionMore work news

Keep reading

Related stories

More from Work