Alibaba's Qwen Audio 3.0 TTS Plus tops the competition in the text-to-speech rankings
Alibaba's Qwen Audio 3.0 just topped the Speech Arena leaderboard. But there's a catch: it's significantly slower than competitors.

Why it matters
Alibaba's multimodal capability expansion positions it as a serious contender in the AI model race, but performance trade-offs (speed vs. quality) reveal ongoing engineering challenges in generalist model development.
The key facts
4 to knowQwen Audio 3.0 TTS Plus ranks #1 on Artificial Analysis' Speech Arena leaderboard
Supports 16 languages with controllable speaking style (natural language + tag-based control like [angry])
Processing speed: 16 characters per second — substantially slower than Sonic 3.5 and Simba 3.2 competitors
Multimodal capability expansion signals Alibaba's push into speech synthesis as core model feature
Go to the source
The Decoderthe-decoder.com
Publisher excerpt: Alibaba's Qwen Audio 3.0 TTS Plus tops Artificial Analysis' Speech Arena leaderboard. The model supports 16 languages and lets users control speaking style with natural language or tags like [angry]. At 16 characters per second, though, it's far slower than rivals Sonic 3.5 and Simba 3.2.