ProgramBench: Can Language Models Rebuild Programs from Scratch?
Can LLMs actually rebuild software from scratch? New benchmark raises uncomfortable questions about AI coding claims.

Why it matters
Academic research benchmarking a critical capability claim in AI — whether language models can truly reconstruct programs — directly relevant to evaluating real-world deployment readiness of AI coding tools that founders and CTOs are betting on.
The key facts
9 to knowProgramBench — new benchmark for evaluating LLM program reconstruction
Tests whether language models can rebuild programs from scratch
Published on arxiv.org (peer-review pending or recent publication)
Low engagement on HN (6 points, 1 comment) suggests niche but credible research audience
Directly challenges coding AI capability claims commonly cited in product/vendor marketing
ProgramBench: new benchmark for evaluating LLM program reconstruction
Published on arXiv (May 7, 2026)
Addresses LLM limitations in program synthesis and code generation
Minimal discussion (6 points, 1 comment on HN) suggests emerging/niche research
Go to the source
Hacker Newsarxiv.org
Publisher excerpt: Article URL: Comments URL: Points: 6 # Comments: 1

