WorkThe story, in brief

Microsoft trained its MAI models on unlicensed web data despite promising "enterprise grade, clean and commercially licensed data"

Microsoft promised 'commercially licensed data.' It trained MAI on unlicensed web crawls anyway.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

Microsoft's training practices contradict its enterprise positioning and clean-data claims, exposing a gap between marketing narrative and actual compliance approach—a critical governance issue for enterprises evaluating AI vendors.

The key facts

5 to know
  1. Microsoft trained MAI models partly on unlicensed web data (Common Crawl)

  2. Marketing claimed 'enterprise grade, clean and commercially licensed data'

  3. Microsoft relies on fair use defense like other AI labs

  4. Training burden placed on site owners to block crawlers

  5. Discrepancy between public claims and actual training practices

Go to the source

The Decoderthe-decoder.com

Publisher excerpt: Microsoft sells its LLM training approach as different from other AI companies. It isn't. The company trained its new MAI models partly on unlicensed web data like Common Crawl, despite claiming they used only "clean and commercially licensed data." Like every other AI lab, Microsoft leans on fair…
Read original report
Back to today's editionMore work news

Keep reading

Related stories

More from Work