Agent skills look great in benchmarks but fall apart under realistic conditions, researchers find
34,000 skills tested. Almost none of them work in the real world. Your agent strategy needs a rethink.

Why it matters
New research reveals a critical gap between benchmark performance and real-world deployment of AI agent skills—a finding that directly challenges current assumptions about modular knowledge architectures and could reshape how companies approach agent capability stacking.
The key facts
4 to knowStudy tested 34,000 real-world AI agent skills
Skills show strong benchmark performance but fail under realistic conditions
Weaker models perform worse with skills than without them
Gap between controlled evaluation and production deployment identified
Go to the source
The Decoderthe-decoder.com
Publisher excerpt: AI agents are supposed to tap into specialized knowledge through so-called skills, modular instructions they can pull up on the fly. But a study testing 34,000 real-world skills finds these enhancements barely help under realistic conditions. Weaker models actually perform worse with them than…