Google AI Releases Gemini 3.5 Transcribe: A Speech-to-Text Model Reporting 2.6% Average WER Across 85+ Languages
Google ships a major speech-to-text capability upgrade with measurable benchmarks (WER, latency, cost) and a deliberate architectural choice (dual endpoints) that practitioners building voice agents need to evaluate against existing transcription stacks.
Why it ranks · · Gemini 3.5 Transcribe released as two endpoints: streaming (sub-second, 4.0% WER, no diarization/timestamps) and batch (2.6% WER, includes diarization/timestamps, 50% lower cost) · 2026-08-28
Read full story