PaperBench: Evaluating AI’s Ability to Replicate AI Research
OpenAI just benchmarked whether AI can replicate AI research. Here's what it found.

Why it matters
PaperBench measures a critical frontier capability: can AI agents autonomously reproduce state-of-the-art research? This directly tests whether AI can accelerate AI development itself—a key inflection point for R&D productivity and competitive moat-building.
The key facts
4 to knowNew benchmark: PaperBench evaluates AI agents on replicating SOTA AI research
Published by OpenAI Apr 2, 2025
Tests agent autonomy in research reproduction—a proxy for AI-driven R&D acceleration
Capability area: research replication and scientific agent performance
Go to the source
OpenAI Blogopenai.com
Publisher excerpt: We introduce PaperBench, a benchmark evaluating the ability of AI agents to replicate state-of-the-art AI research.