OpenWALDO aims to blow the doors off proprietary AI training models
167B tokens. That's OpenWALDO's bet to compete with proprietary giants' trillion-token models.

Why it matters
OpenWALDO is launching a transparent, community-driven training dataset at scale—an attempt to democratize frontier model development by publishing the exact data and methodology behind training. It's a direct challenge to the secrecy and proprietary moats of OpenAI, Anthropic, and Meta.
The key facts
10 to knowOpenWALDO dataset: 167B transparent tokens
Positioned to compete against proprietary models trained on trillions of tokens
Open-weight / transparent methodology model training initiative
Community contribution model for dataset expansion
Challenges proprietary AI training secrecy
OpenWALDO has published 167B transparent tokens
Initiative emphasizes full dataset and training methodology transparency
Positions against proprietary AI training approaches of major labs
Scale gap: 167B tokens vs. trillions deployed by frontier labs
Contributors actively being recruited for the project
Go to the source
The Register AI/MLtheregister.com
Publisher excerpt: Contributors wanted: 167B transparent tokens have a long way to go against AI giants' trillions