PolyAI Releases Dialog-RSN-1: An Audio-Native Dialog Model That Fuses Turn-Taking, Speech Recognition, Function Calling, And Response
A new model architecture that handles end-to-end voice dialog without ASR transcription step, with production latency below the perceptibility threshold. This is a genuine capability advance for voice agents, and practitioners building voice products should track whether this architecture becomes table-stakes.
Why it ranks · Audio-native model architecture (raw audio → response, skipping ASR intermediate) · 2026-07-31
Read full story