Ex-OpenAI researcher bets $100 billion will flow into training data because scaling alone won't cut it
$100B. That's what one ex-OpenAI researcher says the industry will need to spend on training data as scaling hits a wall.

Why it matters
A frontier researcher and OpenAI veteran is betting that model improvement has shifted from compute scale to data quality—a bet that could reshape how labs allocate capital and which startups matter in the next phase of AI development.
The key facts
5 to knowFormer OpenAI researcher Andrew Ho founding a company focused on specialized training data
Models showing specialization trade-offs: improving at coding/math while stagnating or regressing in other domains
Prediction: $100B+ will flow into targeted data collection as scaling approaches limits
Cambridge researcher Adam Hunt co-identifying the problem
Scaling alone insufficient for continued model generalization
Go to the source
The Decoderthe-decoder.com
Publisher excerpt: Former OpenAI employee Andrew Ho and Cambridge researcher Adam Hunt see a growing problem with large language models. Instead of becoming more versatile, the models are becoming more specialized, excelling at coding and math while stagnating or even regressing in other areas. Ho is leaving OpenAI…