FrontierThe story, in brief

Microsoft AI releases new transcription and text-to-speech models for voice agents

Microsoft ships MAI-Transcribe-2-Streaming for real-time voice agents — new transcription and text-to-speech models.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

New speech models for agent infrastructure. Practitioners building voice agents now have access to updated transcription and synthesis capability; the real question is latency, accuracy on accent/noise, and whether this closes gaps versus OpenAI or Anthropic voice products.

The key facts

6 to know
  1. Model: MAI-Transcribe-2-Streaming (real-time transcription)

  2. Category: Text-to-speech and transcription models for voice agents

  3. Vendor: Microsoft AI

  4. Status: Released (not preview)

  5. Measurement: No latency, accuracy, or pricing data disclosed in article

  6. Use case: Voice agent infrastructure

The story so far

Earlier coverage of this storyline

  1. Gemini 3.8 text-to-speech refines voice AI capabilitiesAI Business
  2. Open TTS Leaderboard: Scalable Evaluation for Multilingual Text-to-Speech and Voice CloningHugging Face Blog
  3. Microsoft targets ultra-realistic voice agents with its first streaming transcription modelSiliconAngle
  4. This story

Go to the source

The Decoderthe-decoder.com

Publisher excerpt: Microsoft AI has released MAI-Transcribe-2-Streaming, a new model for real-time transcription.
Read original report
Back to today's editionMore frontier news

Keep reading

Related stories

More from Frontier