The Agent RaceJune 26, 2026via The Decoder

An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run

Why it matters

Epoch AI's MirrorCode benchmark reveals a critical gap in AI reasoning: models can solve routine code reconstruction tasks but hit a wall on complex programs, exposing the limits of current architectures on long-horizon reasoning challenges.

Key signals

  • Claude Opus 4.7 leads MirrorCode benchmark at 56% solve rate
  • Claude rebuilt 16,000-line toolkit in 14 hours
  • Single complex task required 19 days of continuous inference
  • Single task cost $2,600 to run
  • All tested models fail on most complex tasks
  • Benchmark tests code recreation without access to original source

The hook

56%. That's Claude Opus 4.7's solve rate on Epoch AI's new MirrorCode benchmark—but one task took 19 days and $2,600 to crack.

Epoch AI's new MirrorCode benchmark tests whether AI models can recreate complete programs without access to the original code. Claude Opus 4.7 leads with a 56 percent solve rate, rebuilding a 16,000-line toolkit in just 14 hours. But every model tested still fails on the most complex tasks.

The week's key stories, every Friday.

For practitioners and enthusiasts — free, in your inbox.

Free forever. No spam.