Microsoft AI releases new transcription and text-to-speech models for voice agents
Microsoft ships MAI-Transcribe-2-Streaming for real-time voice agents — new transcription and text-to-speech models.

Why it matters
New speech models for agent infrastructure. Practitioners building voice agents now have access to updated transcription and synthesis capability; the real question is latency, accuracy on accent/noise, and whether this closes gaps versus OpenAI or Anthropic voice products.
The key facts
6 to knowModel: MAI-Transcribe-2-Streaming (real-time transcription)
Category: Text-to-speech and transcription models for voice agents
Vendor: Microsoft AI
Status: Released (not preview)
Measurement: No latency, accuracy, or pricing data disclosed in article
Use case: Voice agent infrastructure
The story so far
Earlier coverage of this storyline
Go to the source
The Decoderthe-decoder.com
Publisher excerpt: Microsoft AI has released MAI-Transcribe-2-Streaming, a new model for real-time transcription.