FrontierThe story, in brief

An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run

56%. That's Claude Opus 4.7's solve rate on Epoch AI's new MirrorCode benchmark—but one task took 19 days and $2,600 to crack.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

Epoch AI's MirrorCode benchmark reveals a critical gap in AI reasoning: models can solve routine code reconstruction tasks but hit a wall on complex programs, exposing the limits of current architectures on long-horizon reasoning challenges.

The key facts

6 to know
  1. Claude Opus 4.7 leads MirrorCode benchmark at 56% solve rate

  2. Claude rebuilt 16,000-line toolkit in 14 hours

  3. Single complex task required 19 days of continuous inference

  4. Single task cost $2,600 to run

  5. All tested models fail on most complex tasks

  6. Benchmark tests code recreation without access to original source

Go to the source

The Decoderthe-decoder.com

Publisher excerpt: Epoch AI's new MirrorCode benchmark tests whether AI models can recreate complete programs without access to the original code. Claude Opus 4.7 leads with a 56 percent solve rate, rebuilding a 16,000-line toolkit in just 14 hours. But every model tested still fails on the most complex tasks.
Read original report
Back to today's editionMore frontier news

Keep reading

Related stories

More from Frontier