ElevenLabs' new v4 speech model makes AI voices more expressive and consistent
ElevenLabs v4 cuts latency to 150ms for real-time voice agents—and ranks ahead of Google and Cartesia on the Voice Arena benchmark.

Why it matters
ElevenLabs' new speech model improves expressiveness (laughter, whispers) and consistency for long-form audio, with a Turbo variant enabling real-time agent voice. The release is a meaningful product capability jump in a crowded voice-AI space, but the practical deployment impact for enterprise voice agents remains unclear.
The key facts
5 to knowEleven v4 Turbo latency: 150 milliseconds for speech start
Improved handling: laughter, whispering, voice consistency across long productions
Ranking: v4 ranks ahead of Cartesia and Google Gemini on Artificial Analysis' Voice Arena leaderboard
Use case: audiobooks, real-time voice agents
Status: GA (implied from framing, not explicitly stated)
Go to the source
The Decoderthe-decoder.com
Publisher excerpt: Elevenlabs' new speech model, Eleven v4, follows cues for laughter and whispering more accurately and keeps voices consistent across long productions like audiobooks. Its Turbo variant starts speaking in 150 milliseconds and is built for real-time voice agents. On Artificial Analysis' Voice Arena…