The Atlantic created a searchable database of the music used to train AI
Google and Stability trained on millions of songs without permission. The Atlantic just made it searchable.

Why it matters
Training data transparency and IP accountability are becoming a flashpoint in AI governance. This database exposes the scale of unlicensed content use by major labs, raising questions about artist consent, copyright enforcement, and regulatory liability.
The key facts
5 to knowFour datasets identified: two with 12M and 9M tracks respectively, two with 100K+ tracks each
Google and Stability confirmed use of datasets in research papers
Free Music Archive dataset used for personal streaming but repurposed for AI training
Database made fully searchable and publicly accessible by The Atlantic
Datasets downloaded thousands of times; actual user list unknown
Go to the source
The Verge AItheverge.com
Publisher excerpt: Atlantic reporter Alex Reisner recently uncovered four datasets of music being used to train AI models and made them fully searchable for the public. Two of the sets are absolutely enormous at 12 million and 9 million tracks. The other two are much smaller, but still represent a significant amount…