Alibaba Qwen Releases Qwen-Audio-3.1-Realtime: A Full-Duplex Voice Model Trained to Think, Act, and Decide When to Speak
Alibaba's Qwen-Audio-3.1-Realtime cuts false positives to 13% and hits 82% task success—full-duplex voice reasoning now in production API.

Why it matters
A real-time voice model with reasoning and tool-calling capability, trained to know when NOT to respond, moves from lab to deployed API. Practitioners using voice agents need to benchmark against this; enthusiasts tracking the voice-reasoning frontier should track Alibaba's momentum.
The key facts
5 to knowQwen-Audio-3.1-Realtime released as production API on QwenCloud
Task success: 82.0% (up from 78.4% on τ-Voice adaptation)
Background speech false positives: 13% (down from 73%)
Full-duplex architecture with reasoning, tool-calling, and speech-turn prediction
Available now via API (GA status implied, no preview/beta notation)
The story so far
Earlier coverage of this storyline
Go to the source
MarkTechPostmarktechpost.com
Publisher excerpt: Alibaba's Qwen team released Qwen-Audio-3.1-Realtime, a full-duplex voice model trained to reason, call tools and decide when to speak. On a τ-Voice adaptation, task success rises to 82.0% from 78.4%. Replies to background speech drop from 73% to 13%. It is available now as an API on QwenCloud.