ToolsThe story, in brief

Alibaba’s Tongyi Lab Releases Qwen-Audio-3.0-TTS, a Hosted Text-to-Speech Model in Flash and Plus Tiers Across 16 Languages

Alibaba just shipped Qwen-Audio-3.0-TTS across 16 languages. Two tiers. Hosted only. Here's why that matters for your stack.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

Alibaba is positioning Qwen-Audio as a production-grade multimodal alternative to OpenAI and Google's TTS offerings, expanding the competitive landscape for enterprise speech synthesis. The hosted-only model strategy signals confidence in cloud lock-in but limits developer flexibility.

The key facts

5 to know
  1. Qwen-Audio-3.0-TTS released in two variants: Flash (real-time) and Plus (high-quality)

  2. 16 languages supported

  3. Hosted via Alibaba Cloud Model Studio (not open-source weights)

  4. Production-oriented system designed for four key developer pain points

  5. Multimodal audio capability within Qwen lineage

Go to the source

MarkTechPostmarktechpost.com

Publisher excerpt: Alibaba’s Tongyi Lab has released Qwen-Audio-3.0-TTS, a production-oriented text-to-speech (TTS) system. The model ships in two variants from the same lineage. Flash targets real-time interaction. Plus targets high-quality generation. Both are delivered as hosted models through Alibaba Cloud Model…
Read original report
Back to today's editionMore tools news

Keep reading

Related stories

More from Tools