WorkThe story, in brief

Why Major News Sites Are Blocking The Internet Archive’s Wayback Machine

The Internet Archive just lost access to three decades of news. Here's why publishers are betting against digital preservation.

Illustration of two anonymous hands arranging task cards around an amber tool on a shared desk.
People, judgement and the changing nature of work.AI illustration by KeyNews
The KeyNews take

Why it matters

Major news outlets are blocking the Wayback Machine to combat AI training data scraping, raising critical questions about data access, AI governance, and the long-term preservation of digital history — a policy battle that will shape how AI companies source training data.

The key facts

9 to know
  1. News publishers blocking Wayback Machine access to prevent AI scraper training

  2. Three decades of digital history now restricted

  3. Tension between AI data acquisition and digital preservation

  4. Policy/governance implications for AI training data sourcing

  5. Major news outlets blocking Wayback Machine access

  6. Motivation: preventing AI scraping for training data

  7. Impact on 30+ years of archived digital content

  8. Suggests industry-wide trend toward AI data source restrictions

  9. Raises policy/ethics questions about digital preservation vs. IP protection

Go to the source

Forbes Innovationforbes.com

Publisher excerpt: Major news outlets are blocking the Wayback Machine to fight AI scrapers — and taking three decades of digital history with them.
Read original report
Back to today's editionMore work news

Keep reading

Related stories

More from Work