FrontierAugust 27, 2026via The Decoder
Google's Gemini 3.5 Transcribe turns speech to text in 85 languages while auto-correcting your verbal stumbles
Why it matters
Google shipped a major speech-recognition model with measurably better capability (WER, latency, language coverage, real-time correction) than its predecessor. This is a frontier capability release that sets new bars for multimodal model performance and changes what practitioners expect from speech-to-text systems.
Key signals
- 85 languages supported
- 4.0% word error rate in streaming mode
- 70% lower latency than Chirp 3 predecessor
- Real-time filler-word removal and diarization
- Function calling enables handoff to other Gemini models
- Model name: Gemini 3.5 Transcribe
The hook
4.0% WER, 85 languages, 70% lower latency: Google's Gemini 3.5 Transcribe rewrites the speech-recognition benchmark.
Google's new Gemini 3.5 Transcribe recognizes over 85 languages, strips filler words, and corrects slips of the tongue in real time. It hits a 4.0 percent word error rate in streaming mode, with 70 percent lower latency than its predecessor, Chirp 3. Through function calling, the model can hand off …