AI coding agents can modernize research software but can't judge if the science is right
60x speedups. Zero scientific validation. OpenAI's coding agents modernize research software—but experts warn they're 'eloquent, convincing, and confidently wrong.'

Why it matters
Coding agents show real productivity gains in a real domain (research software modernization), but expose a critical failure mode: agents can optimize for code quality while producing scientifically incorrect results. This shifts the bottleneck from engineering to domain verification—a key operational lesson for deploying agents in expert domains.
The key facts
6 to knowOpenAI and academic partners field report on coding agents modernizing research software
Up to 60x speedups observed
Agents described as 'eloquent, convincing, and confidently wrong in ways that are easy to miss'
Work shifts from code writing to verification of scientific correctness
Deployment context: research software (typically neglected, legacy codebases)
Core finding: agents can pass code metrics while failing domain logic
Go to the source
The Decoderthe-decoder.com
Publisher excerpt: A field report from OpenAI and academic partners shows coding agents can modernize neglected research software, with speedups of up to 60x. But the systems are "eloquent, convincing, and confidently wrong in ways that are easy to miss," participants say. The effort shifts from writing code to the…