Microsoft trained its MAI models on unlicensed web data despite promising "enterprise grade, clean and commercially licensed data"
Microsoft promised 'commercially licensed data.' It trained MAI on unlicensed web crawls anyway.

Why it matters
Microsoft's training practices contradict its enterprise positioning and clean-data claims, exposing a gap between marketing narrative and actual compliance approach—a critical governance issue for enterprises evaluating AI vendors.
The key facts
5 to knowMicrosoft trained MAI models partly on unlicensed web data (Common Crawl)
Marketing claimed 'enterprise grade, clean and commercially licensed data'
Microsoft relies on fair use defense like other AI labs
Training burden placed on site owners to block crawlers
Discrepancy between public claims and actual training practices
Go to the source
The Decoderthe-decoder.com
Publisher excerpt: Microsoft sells its LLM training approach as different from other AI companies. It isn't. The company trained its new MAI models partly on unlicensed web data like Common Crawl, despite claiming they used only "clean and commercially licensed data." Like every other AI lab, Microsoft leans on fair…