FrontierAugust 9, 2026via MarkTechPost
NVIDIA Releases NemotronLabs VoiceChat 11B: An Open Full-Duplex Speech-to-Speech Model with ~450 ms Turn-Taking and Live Tool Calling
Why it matters
A capable open-weight speech model with near-real-time latency and tool integration raises the bar for voice AI capabilities accessible outside closed ecosystems. Practitioners can now evaluate full-duplex voice for production without vendor lock-in.
Key signals
- NemotronLabs VoiceChat 11B: open full-duplex speech-to-speech model
- 448 ms turn-taking latency (near-human conversation threshold)
- 11B parameters (efficient size for deployment)
- Live tool calling capability
- Open weights release
The hook
NVIDIA's 11B speech-to-speech model hits 448ms turn-taking—approaching human conversation speed in open weights.
NVIDIA releases NemotronLabs VoiceChat 11B, an open full-duplex speech-to-speech model with 448 ms latency and live tool calling.
The post NVIDIA Releases NemotronLabs VoiceChat 11B: An Open Full-Duplex Speech-to-Speech Model with ~450 ms Turn-Taking and Live Tool Calling appeared first on MarkTechP…