FrontierAugust 18, 2026via MarkTechPost

Cartesia Ships Sonic-3.6: A Streaming TTS Model That Now Leads Both Artificial Analysis Speech Arenas

Why it matters

A new TTS model architecture and capability benchmark win signals a shift in how speech synthesis is being built. Practitioners evaluating TTS for latency-sensitive applications now have a new reference point; enthusiasts tracking the lab race see state-space models encroaching on transformer dominance outside LLMs.

Key signals

  • Sonic-3.6 ranks #1 on both Artificial Analysis speech leaderboards
  • 1,283 Elo on Provider Voice leaderboard
  • 1,123 Elo on Controlled Voice leaderboard (normalized reference voice synthesis)
  • Sub-90ms time-to-first-audio latency
  • Built on state-space models, not transformers
  • Available in beta on Cartesia API
  • Released August 18, 2026

The hook

Cartesia's Sonic-3.6 just topped both Artificial Analysis speech benchmarks—1,283 Elo on Provider Voice, sub-90ms latency, state-space architecture instead of transformers.

Cartesia has released Sonic-3.6, a streaming text-to-speech model built on state space models rather than transformers. It now ranks #1 on both Artificial Analysis speech leaderboards — 1,283 Elo on Provider Voice and 1,123 on Controlled Voice, the board that clones every model onto the same eight r

The week's key stories, every Friday.

For practitioners and enthusiasts — free, in your inbox.

Free forever. No spam.

Cartesia Ships Sonic-3.6: A Streaming TTS Model That Now Leads Both Artificial Analysis Speech Arenas | KeyNews.AI