The Agent RaceJune 26, 2026via The Decoder
An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run
Why it matters
Epoch AI's MirrorCode benchmark reveals a critical gap in AI reasoning: models can solve routine code reconstruction tasks but hit a wall on complex programs, exposing the limits of current architectures on long-horizon reasoning challenges.
Key signals
- Claude Opus 4.7 leads MirrorCode benchmark at 56% solve rate
- Claude rebuilt 16,000-line toolkit in 14 hours
- Single complex task required 19 days of continuous inference
- Single task cost $2,600 to run
- All tested models fail on most complex tasks
- Benchmark tests code recreation without access to original source
The hook
56%. That's Claude Opus 4.7's solve rate on Epoch AI's new MirrorCode benchmark—but one task took 19 days and $2,600 to crack.
Epoch AI's new MirrorCode benchmark tests whether AI models can recreate complete programs without access to the original code. Claude Opus 4.7 leads with a 56 percent solve rate, rebuilding a 16,000-line toolkit in just 14 hours. But every model tested still fails on the most complex tasks.