What We Learned by Reproducing 2,200 papers from ICML
Hugging Face reproduced 2,200 ICML papers. Here's what broke—and what actually matters.

Why it matters
A large-scale reproducibility audit of frontier research reveals which AI capabilities claims hold up under independent scrutiny, and which don't—directly shaping what practitioners should trust when evaluating next models.
The key facts
5 to know2,200 ICML papers reproduced independently
Hugging Face reproducibility project
August 2026 publication date
Targets capability and benchmark claims in published research
Informs which research results are reliable for downstream practitioners
Go to the source
Hugging Face Bloghuggingface.co