Microsoft takes on AI rivals with three new foundational models
Microsoft just dropped three foundational models. Voice transcription, audio generation, and image creation—all aimed directly at OpenAI and Google.

Why it matters
Microsoft is escalating the AI model wars with multimodal capabilities that directly challenge existing players, signaling intensified competition in foundational AI infrastructure that enterprises rely on for core workflows.
The key facts
5 to knowThree new foundational models released
Voice-to-text transcription capability
Audio generation functionality
Image generation capability
MAI group formed six months ago
Go to the source
TechCrunch AItechcrunch.com
Publisher excerpt: MAI released models that can transcribe voice into text as well as generate audio and images after the group's formation six months ago.