Cartesia Ships Sonic-3.6: A Streaming TTS Model That Now Leads Both Artificial Analysis Speech Arenas
Cartesia's Sonic-3.6 just topped both Artificial Analysis speech benchmarks—1,283 Elo on Provider Voice, sub-90ms latency, state-space architecture instead of transformers.

Why it matters
A new TTS model architecture and capability benchmark win signals a shift in how speech synthesis is being built. Practitioners evaluating TTS for latency-sensitive applications now have a new reference point; enthusiasts tracking the lab race see state-space models encroaching on transformer dominance outside LLMs.
The key facts
7 to knowSonic-3.6 ranks #1 on both Artificial Analysis speech leaderboards
1,283 Elo on Provider Voice leaderboard
1,123 Elo on Controlled Voice leaderboard (normalized reference voice synthesis)
Sub-90ms time-to-first-audio latency
Built on state-space models, not transformers
Available in beta on Cartesia API
Released August 18, 2026
Go to the source
MarkTechPostmarktechpost.com
Publisher excerpt: Cartesia has released Sonic-3.6, a streaming text-to-speech model built on state space models rather than transformers. It now ranks #1 on both Artificial Analysis speech leaderboards — 1,283 Elo on Provider Voice and 1,123 on Controlled Voice, the board that clones every model onto the same eight…