Microsoft AI Releases MAI-Transcribe-2-Streaming: #1 Real-Time Speech-to-Text Model on Artificial Analysis
#1 on Artificial Analysis. Microsoft's MAI-Transcribe-2-Streaming hits 2.5% WER in real-time across 60 languages.

Why it matters
Microsoft ships a production speech-to-text model ranked atop streaming benchmarks with measured latency and multilingual support. For practitioners building voice agents and conversational AI, this is a concrete alternative to existing transcription APIs with published performance data — but public preview status and introductory pricing mean watch-and-wait before migrating production workloads.
The key facts
9 to knowMAI-Transcribe-2-Streaming ranks #1 of 38 models on Artificial Analysis AA-WER Streaming benchmark
2.5% WER at 0.13s latency on final transcripts; 2.5% WER at 0.12s on first partials
Covers 60 languages with continuous language detection
Introductory pricing: $0.54 per hour
Available now in public preview on Microsoft Foundry
Source: MarkTechPost reporting on Microsoft AI announcement
2.5% WER (word error rate) at 0.13s latency on final transcripts; 2.5% WER at 0.12s on first partials
Status: Public preview on Microsoft Foundry
Source: MarkTechPost (no independent testing reported)
The story so far
Earlier coverage of this storyline
- Microsoft AI releases new transcription and text-to-speech models for voice agentsThe Decoder
- This story
Go to the source
MarkTechPostmarktechpost.com
Publisher excerpt: Microsoft AI has released MAI-Transcribe-2-Streaming, its first real-time speech-to-text model. It ranks #1 of 38 models on Artificial Analysis AA-WER Streaming. It scores 2.5% WER at 0.13s on final transcripts and 2.5% at 0.12s on first partials. It covers 60 languages with continuous language…