FrontierThe story, in brief

Import AI 466: The bitter lesson for robotics, AIs complete week-long programming tasks; and OpenAI’s accidental AI hacker

AI systems just completed week-long programming tasks. Here's what that means for your engineering roadmap.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

New benchmark (MirrorCode) from Epoch and METR reveals AI capability limits on long-horizon coding—a critical signal for teams betting on AI-assisted development and autonomous agents.

The key facts

5 to know
  1. Epoch and METR released MirrorCode benchmark for long-horizon programming tasks

  2. AI systems cannot yet solve hardest benchmark tasks

  3. Week-long programming task completion reported

  4. Published July 27, 2026

  5. Newsletter format limits granular data extraction

Go to the source

Import AI (Blog)jack-clark.net

Publisher excerpt: Welcome to Import AI, a newsletter about AI research. Import AI runs on arXiv, cappuccinos, and feedback from readers. If you’d like to support this, please subscribe. Subscribe now Epoch and METR release MirrorCode, a benchmark for seeing how well AI systems can do long-horizon programming…
Read original report
Back to today's editionMore frontier news

Keep reading

Related stories

More from Frontier