FrontierThe story, in brief

Nvidia drops a free 100M-parameter model that identifies up to eight speakers in real time

Nvidia just open-sourced a 100M-parameter speaker-diarization model. It's free, identifies up to eight speakers in real time, and runs locally.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

A capable, open-weight model for speaker identification lowers the barrier for speech-processing applications. Practitioners can now embed diarization without licensing proprietary APIs or building from scratch, but the 8-speaker limit and real-time performance at scale remain unstated.

The key facts

12 to know
  1. Model: Nemotron 3 Diarization

  2. Parameters: 100M

  3. Capability: speaker identification in conversations

  4. Max speakers handled: 8

  5. Availability: free, open-weight release

  6. Real-time inference claimed

  7. Specific latency, accuracy benchmark, or supported languages not disclosed

  8. Integration path and deployment requirements not detailed

  9. Capability: identifies up to 8 speakers in real time

  10. License: open/free

  11. Vendor: Nvidia

  12. Task: speaker diarization (who spoke when)

The story so far

Earlier coverage of this storyline

  1. NVIDIA Releases Nemotron 3 Diarization: A 100M-Parameter Open-Weight Model That Tracks 8 Speakers in Real TimeMarkTechPost
  2. This story

Go to the source

The Decoderthe-decoder.com

Publisher excerpt: Nvidia released Nemotron 3 Diarization, an AI model that identifies which speaker is talking at any given moment in a conversation.
Read original report
Back to today's editionMore frontier news

Keep reading

Related stories

More from Frontier