Study contradicts Anthropic and OpenAI claims that autonomous AI research is within reach
Claude Opus 4.8 and GPT-5.6 Sol can execute research workflows—but can't think like researchers. A new study exposes the gap between autonomy hype and autonomous judgment.

Why it matters
A peer-reviewed evaluation from Princeton and the UK AI Security Institute directly contradicts frontier lab claims about near-term autonomous research capability. The finding reshapes how practitioners should think about agent-native model roadmaps and capability timelines.
The key facts
7 to knowStudy tested Claude Opus 4.8 and GPT-5.6 Sol on independent AI research paper writing
Models given 6 days, $3,000 API credits, GPU access
Original paper authors rated agent outputs as 'Reject'
Models can handle full research engineering but fail on research judgment, creative problem-solving, approach abandonment
Conducted by Princeton and UK AI Security Institute
Published Aug 14, 2026
Directly contradicts Anthropic and OpenAI public claims on autonomous research timeline
Go to the source
The Decoderthe-decoder.com
Publisher excerpt: AI agents using Claude Opus 4.8 and GPT-5.6 Sol were given six days, $3,000 in API credits, and GPU access to independently write AI research papers. The original authors of unpublished NeurIPS papers rated the results as "Reject." According to the study, conducted with Princeton and the UK AI…