FrontierThe story, in brief

Cartesia Ships Sonic-3.6: A Streaming TTS Model That Now Leads Both Artificial Analysis Speech Arenas

Cartesia's Sonic-3.6 just topped both Artificial Analysis speech benchmarks—1,283 Elo on Provider Voice, sub-90ms latency, state-space architecture instead of transformers.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

A new TTS model architecture and capability benchmark win signals a shift in how speech synthesis is being built. Practitioners evaluating TTS for latency-sensitive applications now have a new reference point; enthusiasts tracking the lab race see state-space models encroaching on transformer dominance outside LLMs.

The key facts

7 to know
  1. Sonic-3.6 ranks #1 on both Artificial Analysis speech leaderboards

  2. 1,283 Elo on Provider Voice leaderboard

  3. 1,123 Elo on Controlled Voice leaderboard (normalized reference voice synthesis)

  4. Sub-90ms time-to-first-audio latency

  5. Built on state-space models, not transformers

  6. Available in beta on Cartesia API

  7. Released August 18, 2026

Go to the source

MarkTechPostmarktechpost.com

Publisher excerpt: Cartesia has released Sonic-3.6, a streaming text-to-speech model built on state space models rather than transformers. It now ranks #1 on both Artificial Analysis speech leaderboards — 1,283 Elo on Provider Voice and 1,123 on Controlled Voice, the board that clones every model onto the same eight…
Read original report
Back to today's editionMore frontier news

Keep reading

Related stories

More from Frontier