NVIDIA Releases Nemotron 3 Diarization: A 100M-Parameter Open-Weight Model That Tracks 8 Speakers in Real Time
NVIDIA drops a 100M-parameter speaker diarization model that handles 8 overlapping voices in real time—and it's open-weight.

Why it matters
A capable, deployable open-weight model for a real production need (speaker tracking in transcription, meeting analytics, surveillance) signals NVIDIA's strategy to compete in the foundation-model commons alongside capability releases. Practitioners building audio pipelines now have a licensed alternative to proprietary or academic-only solutions.
The key facts
6 to know100M parameters
Tracks up to 8 speakers including overlapping voices
Supports both offline and real-time streaming from single checkpoint
Open-weight release on Hugging Face
Diarization task: speaker identification in multi-speaker audio
Production-ready deployment capability
The story so far
Earlier coverage of this storyline
Go to the source
MarkTechPostmarktechpost.com
Publisher excerpt: NVIDIA has released Nemotron 3 Diarization, an open-weight speaker diarization model on Hugging Face. It answers one question about any conversation: who spoke when. The 100M-parameter model tracks up to 8 speakers, including when voices overlap. One checkpoint handles both offline recordings and…