An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run
56%. That's Claude Opus 4.7's solve rate on Epoch AI's new MirrorCode benchmark—but one task took 19 days and $2,600 to crack.

Why it matters
Epoch AI's MirrorCode benchmark reveals a critical gap in AI reasoning: models can solve routine code reconstruction tasks but hit a wall on complex programs, exposing the limits of current architectures on long-horizon reasoning challenges.
The key facts
6 to knowClaude Opus 4.7 leads MirrorCode benchmark at 56% solve rate
Claude rebuilt 16,000-line toolkit in 14 hours
Single complex task required 19 days of continuous inference
Single task cost $2,600 to run
All tested models fail on most complex tasks
Benchmark tests code recreation without access to original source
Go to the source
The Decoderthe-decoder.com
Publisher excerpt: Epoch AI's new MirrorCode benchmark tests whether AI models can recreate complete programs without access to the original code. Claude Opus 4.7 leads with a 56 percent solve rate, rebuilding a 16,000-line toolkit in just 14 hours. But every model tested still fails on the most complex tasks.