ToolsThe story, in brief

Google's new Flash TTS models let you design AI voices from scratch using text descriptions

Google just shipped text-to-speech that builds voices from descriptions—no cloning needed. Over 100 languages, stage directions, two-voice dialogue from one script.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

Google's new Flash TTS models expand the accessibility and creative control of AI voice generation for developers and creators, moving beyond simple speech synthesis to scriptable, multi-voice dialogue production at scale.

The key facts

10 to know
  1. Gemini 3.8 Flash TTS and Flash-Lite TTS announced

  2. Supports 100+ languages

  3. Create new voices from text descriptions (generative voice design)

  4. Voice cloning from 30-second audio sample

  5. Stage directions per line support

  6. Two-voice dialogue generation from single script

  7. Two new models: Gemini 3.8 Flash TTS and Flash-Lite TTS

  8. Support for 100+ languages

  9. Voice design from text descriptions (e.g., 'warm, authoritative female voice')

  10. Stage directions for individual lines

Go to the source

The Decoderthe-decoder.com

Publisher excerpt: Google is introducing two new text-to-speech models, Gemini 3.8 Flash TTS and Flash-Lite TTS, which support more than 100 languages. Flash TTS can create new voices from text descriptions, and both models let users add stage directions to individual lines and generate two-voice dialogue from a…
Read original report
Back to today's editionMore tools news

Keep reading

Related stories

More from Tools