Import AI 466: The bitter lesson for robotics, AIs complete week-long programming tasks; and OpenAI’s accidental AI hacker
AI systems just completed week-long programming tasks. Here's what that means for your engineering roadmap.

Why it matters
New benchmark (MirrorCode) from Epoch and METR reveals AI capability limits on long-horizon coding—a critical signal for teams betting on AI-assisted development and autonomous agents.
The key facts
5 to knowEpoch and METR released MirrorCode benchmark for long-horizon programming tasks
AI systems cannot yet solve hardest benchmark tasks
Week-long programming task completion reported
Published July 27, 2026
Newsletter format limits granular data extraction
Go to the source
Import AI (Blog)jack-clark.net
Publisher excerpt: Welcome to Import AI, a newsletter about AI research. Import AI runs on arXiv, cappuccinos, and feedback from readers. If you’d like to support this, please subscribe. Subscribe now Epoch and METR release MirrorCode, a benchmark for seeing how well AI systems can do long-horizon programming…