FrontierThe story, in brief

Alibaba Qwen Releases Qwen-Audio-3.1-Realtime: A Full-Duplex Voice Model Trained to Think, Act, and Decide When to Speak

Alibaba's Qwen-Audio-3.1-Realtime cuts false positives to 13% and hits 82% task success—full-duplex voice reasoning now in production API.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

A real-time voice model with reasoning and tool-calling capability, trained to know when NOT to respond, moves from lab to deployed API. Practitioners using voice agents need to benchmark against this; enthusiasts tracking the voice-reasoning frontier should track Alibaba's momentum.

The key facts

5 to know
  1. Qwen-Audio-3.1-Realtime released as production API on QwenCloud

  2. Task success: 82.0% (up from 78.4% on τ-Voice adaptation)

  3. Background speech false positives: 13% (down from 73%)

  4. Full-duplex architecture with reasoning, tool-calling, and speech-turn prediction

  5. Available now via API (GA status implied, no preview/beta notation)

The story so far

Earlier coverage of this storyline

  1. Alibaba launches Qwen Audio 3.1 with new models and slashes AI audio prices by up to 95 percentThe Decoder
  2. This story

Go to the source

MarkTechPostmarktechpost.com

Publisher excerpt: Alibaba's Qwen team released Qwen-Audio-3.1-Realtime, a full-duplex voice model trained to reason, call tools and decide when to speak. On a τ-Voice adaptation, task success rises to 82.0% from 78.4%. Replies to background speech drop from 73% to 13%. It is available now as an API on QwenCloud.
Read original report
Back to today's editionMore frontier news

Keep reading

Related stories

More from Frontier