Old OCR text cripples language model training, and FineBooks wants to fix that at scale
97.6% accuracy, $2 per 1K pages—Hugging Face and EleutherAI solve the OCR bottleneck starving LLM training on historical texts.

Why it matters
Training-data quality is a hard constraint on frontier models. FineBooks benchmarks OCR at scale and proves cost-effective solutions exist—practitioners sourcing historical text corpora now have a playbook.
The key facts
12 to knowFineBooks tested 14 open-source OCR models
2,000+ historical book pages evaluated
Top performer: dots.mocr at 97.6% character accuracy
Cost: under $2 per 1,000 pages
Suitable for AI training (not yet for scholarly transcription)
Hugging Face + EleutherAI collaboration
FineBooks tested 14 open-source OCR models on 2,000+ historical book pages
Top model (dots.mocr) achieves 97.6% character accuracy
Cost: under $2 per thousand pages
Accuracy sufficient for AI training; not yet for scholarly transcription
Partnership: Hugging Face and EleutherAI
Focus: removing OCR degradation as a training-data constraint
Go to the source
The Decoderthe-decoder.com
Publisher excerpt: The FineBooks project from Hugging Face and EleutherAI tested 14 open-source OCR models on more than 2,000 historical book pages. The top model, dots.mocr, hits 97.6 percent character accuracy at under two dollars per thousand pages. That's good enough for AI training data, but not yet for…