AI agents overstate their results and remain far from autonomous research, study finds
GPT-5.6 Sol hit 15% of human performance on autonomous research. Current AI agents can run experiments—but can't think critically about what went wrong.

Why it matters
Epoch AI and Anthropic's independent findings reveal a critical limitation in agentic AI: models can execute multi-step research workflows but lack self-correction and genuine scientific reasoning. This matters for enterprise deployments betting on autonomous research and analysis—the constraint is cognitive, not just architectural.
The key facts
6 to knowGPT-5.6 Sol reached 15% of human reference score on autonomous research tasks
Claude Fable 5 showed similar limitations
Both models used only known methods, no novel approaches discovered
Primary failure mode: inability to critically question own results
Study by Epoch AI and Anthropic (independent confirmation)
Models can run experiments but lack scientific self-criticism
Go to the source
The Decoderthe-decoder.com
Publisher excerpt: Epoch AI and Anthropic independently found the same thing: current AI models like GPT-5.6 Sol and Claude Fable 5 can run experiments but lack scientific self-criticism and genuine creative thinking. At best, Sol reached 15 percent of the human reference score, and even that came from methods…