WorkThe story, in brief

The Atlantic created a searchable database of the music used to train AI

Google and Stability trained on millions of songs without permission. The Atlantic just made it searchable.

Illustration of two anonymous hands arranging task cards around an amber tool on a shared desk.
People, judgement and the changing nature of work.AI illustration by KeyNews
The KeyNews take

Why it matters

Training data transparency and IP accountability are becoming a flashpoint in AI governance. This database exposes the scale of unlicensed content use by major labs, raising questions about artist consent, copyright enforcement, and regulatory liability.

The key facts

5 to know
  1. Four datasets identified: two with 12M and 9M tracks respectively, two with 100K+ tracks each

  2. Google and Stability confirmed use of datasets in research papers

  3. Free Music Archive dataset used for personal streaming but repurposed for AI training

  4. Database made fully searchable and publicly accessible by The Atlantic

  5. Datasets downloaded thousands of times; actual user list unknown

Go to the source

The Verge AItheverge.com

Publisher excerpt: Atlantic reporter Alex Reisner recently uncovered four datasets of music being used to train AI models and made them fully searchable for the public. Two of the sets are absolutely enormous at 12 million and 9 million tracks. The other two are much smaller, but still represent a significant amount…
Read original report
Back to today's editionMore work news

Keep reading

Related stories

More from Work