AgentsThe story, in brief

AI agents overstate their results and remain far from autonomous research, study finds

GPT-5.6 Sol hit 15% of human performance on autonomous research. Current AI agents can run experiments—but can't think critically about what went wrong.

Illustration of independent geometric mechanisms passing paper tasks along branching amber tracks.
AI agents and the coordination of work.AI illustration by KeyNews
The KeyNews take

Why it matters

Epoch AI and Anthropic's independent findings reveal a critical limitation in agentic AI: models can execute multi-step research workflows but lack self-correction and genuine scientific reasoning. This matters for enterprise deployments betting on autonomous research and analysis—the constraint is cognitive, not just architectural.

The key facts

6 to know
  1. GPT-5.6 Sol reached 15% of human reference score on autonomous research tasks

  2. Claude Fable 5 showed similar limitations

  3. Both models used only known methods, no novel approaches discovered

  4. Primary failure mode: inability to critically question own results

  5. Study by Epoch AI and Anthropic (independent confirmation)

  6. Models can run experiments but lack scientific self-criticism

Go to the source

The Decoderthe-decoder.com

Publisher excerpt: Epoch AI and Anthropic independently found the same thing: current AI models like GPT-5.6 Sol and Claude Fable 5 can run experiments but lack scientific self-criticism and genuine creative thinking. At best, Sol reached 15 percent of the human reference score, and even that came from methods…
Read original report
Back to today's editionMore agents news

Keep reading

Related stories

More from Agents