FrontierThe story, in brief

Alibaba's Qwen Audio 3.0 TTS Plus tops the competition in the text-to-speech rankings

Alibaba's Qwen Audio 3.0 just topped the Speech Arena leaderboard. But there's a catch: it's significantly slower than competitors.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

Alibaba's multimodal capability expansion positions it as a serious contender in the AI model race, but performance trade-offs (speed vs. quality) reveal ongoing engineering challenges in generalist model development.

The key facts

4 to know
  1. Qwen Audio 3.0 TTS Plus ranks #1 on Artificial Analysis' Speech Arena leaderboard

  2. Supports 16 languages with controllable speaking style (natural language + tag-based control like [angry])

  3. Processing speed: 16 characters per second — substantially slower than Sonic 3.5 and Simba 3.2 competitors

  4. Multimodal capability expansion signals Alibaba's push into speech synthesis as core model feature

Go to the source

The Decoderthe-decoder.com

Publisher excerpt: Alibaba's Qwen Audio 3.0 TTS Plus tops Artificial Analysis' Speech Arena leaderboard. The model supports 16 languages and lets users control speaking style with natural language or tags like [angry]. At 16 characters per second, though, it's far slower than rivals Sonic 3.5 and Simba 3.2.
Read original report
Back to today's editionMore frontier news

Keep reading

Related stories

More from Frontier