← All posts / Models

Eight Voices, One Checkpoint: NVIDIA's Nemotron 3 Diarization Rewrites Who-Spoke-When

NVIDIA's new open-weight 100M-parameter model tracks up to 8 overlapping speakers in real time, cuts diarization error by a third on DIHARD III, and tops VoiceArena's new Diarization-Bench at 14.72% DER.

Eight Voices, One Checkpoint: NVIDIA's Nemotron 3 Diarization Rewrites Who-Spoke-When

Every voice AI engineer eventually hits the same wall: the transcript is perfect, the timing is perfect, and the output is still useless — because nobody knows who said what. Speaker diarization, the unglamorous task of answering “who spoke when,” is the load-bearing wall of meeting transcription, call-center analytics, and any voice agent that has to survive a real conversation with interruptions and crosstalk. On September 23, NVIDIA took a serious swing at that wall with an open-weight model that handles twice as many speakers as its predecessor at a fraction of the size.

Nemotron 3 Diarization is a roughly 100M-parameter speaker diarization model released under NVIDIA’s open Model & Weights license (OpenMDW 1.1), which explicitly permits commercial use. It determines “who spoke when” in real-world audio, supports both streaming and offline inference, and — the headline feature — handles up to eight simultaneous speakers, doubling the four-speaker ceiling of the previous-generation Streaming Sortformer. The release includes a live demo on Hugging Face Spaces pairing the diarizer with streaming ASR, and a new lightweight C++ runtime called NeMo-Speech.cpp that can label a meeting recording with a single command line.

One checkpoint, four latency profiles

The most technically interesting design choice is that a single checkpoint serves every operating point. The same weights run in an offline-style configuration with a 30.4-second input buffer, or in streaming configurations at 1.04 s, 0.64 s, and 0.32 s of input-buffer latency — with 80 ms available for latency-critical applications, though NVIDIA recommends 0.32 s as the practical floor. Output frame resolution is configurable in multiples of 10 ms, and chunked inference removes any limit on audio duration.

This is not a trivial engineering feat. Streaming diarization is hard because speaker identity must persist across chunks: when person three starts talking 40 seconds into a call, the model has to decide whether this is a new voice or speaker two returning. Nemotron 3 inherits its solution from the Streaming Sortformer research line — an Arrival-Order Speaker Cache (AOSC) that retains speaker information from earlier chunks, paired with a FIFO queue carrying recent frame context. The model resolves the classic speaker-permutation ambiguity by ordering its output channels according to each speaker’s first arrival in the audio. Following the Sortformer architecture, that ordering is baked into the model’s behavior rather than bolted on with post-processing.

The numbers

On DIHARD III — the 11-domain benchmark widely regarded as the hardest general diarization test — the gains are large. In offline configuration (30.4 s buffer), Nemotron 3 scores 12.73% full DER versus 19.09% for the previous diar_streaming_sortformer_4spk-v2.1 baseline, a 33% relative reduction. Crucially, the model holds most of that advantage under streaming pressure: at 1.04 s latency it still posts 13.18% DER, and even at the 0.32 s ultra-low-latency setting it stays at 13.55% — both comfortably below the offline score of the old baseline. The hard 5–9-speaker slice of DIHARD improves from 40.21% to 27.58% DER. Speaker counting also sharpens: counting accuracy rises from 75.29% to 81.47%, and mean absolute counting error drops from 0.51 to 0.27.

An important methodological note that NVIDIA itself emphasizes: the reported scores use forced-alignment reference labels for AMI, AliMeeting, and NOTSOFAR1, because the original segment-level annotations were built for transcription rather than frame-accurate diarization and can penalize correct silence predictions. Anyone reproducing these numbers must use the exact published reference RTTMs — different labels mean a different protocol and non-comparable results.

Beyond the academic benchmarks, the model debuted at #1 on VoiceArena’s new Diarization-Bench leaderboard, scoring 14.72% DER across sessions of 2 to 8 speakers and more than 280 unique speakers — roughly a 24% relative margin over the runner-up among 12 ranked systems. Across the eight evaluation conditions NVIDIA lists, the model averages a 41.0% relative DER reduction.

How it was trained

The training recipe is a two-stage pipeline, and it is fully reproducible from open code. Training was initialized from a Transformer-based NEST self-supervised checkpoint, then run on 8 nodes of 8×A100-80GB in two phases: first offline training on simulated multi-speaker data, then streaming fine-tuning on a combination of real conversations and simulated mixtures. The fine-tuning data spans 901 condition-specific recordings — telephonic speech (CALLHOME), far-field and headset meetings (AMI, NOTSOFAR1), Mandarin meetings (AliMeeting), and the multilingual DIHARD III collection. Both the example scripts and the base configs for each stage are public in the NVIDIA-NeMo/Speech repository.

The inference stack is equally accessible. In Python, the model loads through NeMo as a SortformerEncLabelModel; word-level speaker tagging for a transcript is essentially a one-liner on top. The new NeMo-Speech.cpp C++ runtime targets local deployment: nemo-speech diarize meeting.wav produces segments, and nemo-speech transcribe meeting.wav --diarize --json fuses diarization with transcription for word-attributed JSON output. Input is plain 16 kHz mono audio; output is a per-speaker activity grid at 10 ms-class resolution.

Why it matters

The boring parts of the voice stack are where deployment reality lives. Full-duplex speech models like NVIDIA’s own 12B Nemotron 3 VoiceChat handle turn-taking natively, but the far larger market — meeting notes, call-center QA, courtroom and medical transcription, robot ears in multi-person rooms — runs on pipeline architectures where diarization quality caps the ceiling of the entire system. A 100M open-weight model that runs at 0.32 s latency without losing accuracy moves that capability from “cloud API with per-minute billing” to “runs next to your ASR, wherever your ASR runs.” Third parties are already moving: FluidInference has ported the model to CoreML with full Apple Neural Engine coverage, and voice-app builders had on-device variants shipping within a day.

It also continues a deliberate NVIDIA pattern: Parakeet for ASR, Nemotron for text reasoning, now a diarizer — an open-weight answer to every layer of the speech stack, each small enough and permissive enough to seed adoption. The frontier-model conversation obsesses over trillion-parameter reasoning systems, but the models quietly standardizing the industry’s plumbing look a lot more like this one: 100M parameters, one job, four latency profiles, and a license that lets you ship it.