FrontierThe story, in brief

OpenWALDO aims to blow the doors off proprietary AI training models

167B tokens. That's OpenWALDO's bet to compete with proprietary giants' trillion-token models.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

OpenWALDO is launching a transparent, community-driven training dataset at scale—an attempt to democratize frontier model development by publishing the exact data and methodology behind training. It's a direct challenge to the secrecy and proprietary moats of OpenAI, Anthropic, and Meta.

The key facts

10 to know
  1. OpenWALDO dataset: 167B transparent tokens

  2. Positioned to compete against proprietary models trained on trillions of tokens

  3. Open-weight / transparent methodology model training initiative

  4. Community contribution model for dataset expansion

  5. Challenges proprietary AI training secrecy

  6. OpenWALDO has published 167B transparent tokens

  7. Initiative emphasizes full dataset and training methodology transparency

  8. Positions against proprietary AI training approaches of major labs

  9. Scale gap: 167B tokens vs. trillions deployed by frontier labs

  10. Contributors actively being recruited for the project

Go to the source

The Register AI/MLtheregister.com

Publisher excerpt: Contributors wanted: 167B transparent tokens have a long way to go against AI giants' trillions
Read original report
Back to today's editionMore frontier news

Keep reading

Related stories

More from Frontier