FrontierAugust 18, 2026via MarkTechPost
Cartesia Ships Sonic-3.6: A Streaming TTS Model That Now Leads Both Artificial Analysis Speech Arenas
Why it matters
A new TTS model architecture and capability benchmark win signals a shift in how speech synthesis is being built. Practitioners evaluating TTS for latency-sensitive applications now have a new reference point; enthusiasts tracking the lab race see state-space models encroaching on transformer dominance outside LLMs.
Key signals
- Sonic-3.6 ranks #1 on both Artificial Analysis speech leaderboards
- 1,283 Elo on Provider Voice leaderboard
- 1,123 Elo on Controlled Voice leaderboard (normalized reference voice synthesis)
- Sub-90ms time-to-first-audio latency
- Built on state-space models, not transformers
- Available in beta on Cartesia API
- Released August 18, 2026
The hook
Cartesia's Sonic-3.6 just topped both Artificial Analysis speech benchmarks—1,283 Elo on Provider Voice, sub-90ms latency, state-space architecture instead of transformers.
Cartesia has released Sonic-3.6, a streaming text-to-speech model built on state space models rather than transformers. It now ranks #1 on both Artificial Analysis speech leaderboards — 1,283 Elo on Provider Voice and 1,123 on Controlled Voice, the board that clones every model onto the same eight r…