FrontierAugust 27, 2026via The Decoder

Google's Gemini 3.5 Transcribe turns speech to text in 85 languages while auto-correcting your verbal stumbles

Why it matters

Google shipped a major speech-recognition model with measurably better capability (WER, latency, language coverage, real-time correction) than its predecessor. This is a frontier capability release that sets new bars for multimodal model performance and changes what practitioners expect from speech-to-text systems.

Key signals

  • 85 languages supported
  • 4.0% word error rate in streaming mode
  • 70% lower latency than Chirp 3 predecessor
  • Real-time filler-word removal and diarization
  • Function calling enables handoff to other Gemini models
  • Model name: Gemini 3.5 Transcribe

The hook

4.0% WER, 85 languages, 70% lower latency: Google's Gemini 3.5 Transcribe rewrites the speech-recognition benchmark.

Google's new Gemini 3.5 Transcribe recognizes over 85 languages, strips filler words, and corrects slips of the tongue in real time. It hits a 4.0 percent word error rate in streaming mode, with 70 percent lower latency than its predecessor, Chirp 3. Through function calling, the model can hand off

The week's key stories, every Friday.

For practitioners and enthusiasts — free, in your inbox.

Free forever. No spam.