ToolsThe story, in brief

ElevenLabs' new v4 speech model makes AI voices more expressive and consistent

ElevenLabs v4 cuts latency to 150ms for real-time voice agents—and ranks ahead of Google and Cartesia on the Voice Arena benchmark.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

ElevenLabs' new speech model improves expressiveness (laughter, whispers) and consistency for long-form audio, with a Turbo variant enabling real-time agent voice. The release is a meaningful product capability jump in a crowded voice-AI space, but the practical deployment impact for enterprise voice agents remains unclear.

The key facts

5 to know
  1. Eleven v4 Turbo latency: 150 milliseconds for speech start

  2. Improved handling: laughter, whispering, voice consistency across long productions

  3. Ranking: v4 ranks ahead of Cartesia and Google Gemini on Artificial Analysis' Voice Arena leaderboard

  4. Use case: audiobooks, real-time voice agents

  5. Status: GA (implied from framing, not explicitly stated)

Go to the source

The Decoderthe-decoder.com

Publisher excerpt: Elevenlabs' new speech model, Eleven v4, follows cues for laughter and whispering more accurately and keeps voices consistent across long productions like audiobooks. Its Turbo variant starts speaking in 150 milliseconds and is built for real-time voice agents. On Artificial Analysis' Voice Arena…
Read original report
Back to today's editionMore tools news

Keep reading

Related stories

More from Tools