Alibaba launches Qwen Audio 3.1 with new models and slashes AI audio prices by up to 95 percent
Alibaba's Qwen Audio 3.1 cuts speech AI pricing by up to 95%—multilingual ASR, emotion detection, real-time TTS, all shipping this week.

Why it matters
Alibaba is shipping a competitive open audio model lineup with aggressive pricing that undercuts incumbents and lowers the barrier to deploying speech AI in production—a significant move in the models race.
The key facts
6 to knowFive new audio models: ASR, ASR-Next (multi-speaker + emotion + ambient detection), TTS, real-time interaction variants
Up to 95% price reduction on AI audio services
Multilingual and dialect support with automatic filler-word cleanup
Emotion detection, ambient sound classification, machine noise detection
Real-time TTS and interaction capabilities
Qwen (Alibaba's AI lab) entry into frontier audio models—competitive with OpenAI Whisper, Google Speech-to-Text
Go to the source
The Decoderthe-decoder.com
Publisher excerpt: Alibaba's AI team Qwen has released Qwen-Audio-3.1, a lineup of five models for speech recognition (ASR), text-to-speech (TTS), and real-time interaction. The ASR model improves multilingual and dialect recognition and automatically cleans up filler words and repetitions. ASR-Next adds…