Nvidia drops a free 100M-parameter model that identifies up to eight speakers in real time
Nvidia just open-sourced a 100M-parameter speaker-diarization model. It's free, identifies up to eight speakers in real time, and runs locally.

Why it matters
A capable, open-weight model for speaker identification lowers the barrier for speech-processing applications. Practitioners can now embed diarization without licensing proprietary APIs or building from scratch, but the 8-speaker limit and real-time performance at scale remain unstated.
The key facts
12 to knowModel: Nemotron 3 Diarization
Parameters: 100M
Capability: speaker identification in conversations
Max speakers handled: 8
Availability: free, open-weight release
Real-time inference claimed
Specific latency, accuracy benchmark, or supported languages not disclosed
Integration path and deployment requirements not detailed
Capability: identifies up to 8 speakers in real time
License: open/free
Vendor: Nvidia
Task: speaker diarization (who spoke when)
The story so far
Earlier coverage of this storyline
Go to the source
The Decoderthe-decoder.com
Publisher excerpt: Nvidia released Nemotron 3 Diarization, an AI model that identifies which speaker is talking at any given moment in a conversation.