Run Nemotron 3 Diarization for Streaming and Offline Speaker Labels
NVIDIA's 100M-parameter diarizer labels up to 8 speakers, streaming with 0.32 s of buffer or offline with 30.4 s. Here's how to run it.
The code in this post comes from the model's docs and hasn't been run in our CI yet. If something breaks, let us know.
Nemotron 3 Diarization is NVIDIA’s open-weight speaker diarization model. It answers “who spoke when” for live or recorded audio and handles up to eight speakers. It follows the Sortformer design, so its eight output channels are ordered by when each speaker first talks. For streaming it uses the Arrival-Order Speaker Cache (AOSC) and FIFO queue from Streaming Sortformer. The main thing that sets it apart is that one 100M-parameter checkpoint covers both cases. You can run it offline with a 30.4 s input buffer or stream it with a buffer as short as 0.32 s, and audio length has no upper limit when you use chunked inference.
Key specs
| Parameters | 100M |
| Architecture | 31-layer Transformer encoder with RoPE, 80 ms encoder frames, Conv1D upsampling to 10 ms |
| Max speakers | 8 |
| Input | 16 kHz mono audio (.wav, .flac, .opus, .mp3) |
| Output | [T, 8] per-speaker activity probabilities, 10 ms frames by default |
| Latency configs | 30.4 s (offline), 1.04 s, 0.64 s, 0.32 s |
| Training data | about 10,000 h of real conversations plus 82,611 h of simulated mixtures |
| Runtimes | NeMo Speech, Transformers (install from source), NeMo-Speech.cpp |
The card compares the model against NVIDIA’s previous release, diar_streaming_sortformer_4spk-v2.1. These are selected DER results (lower is better):
| Benchmark | v2.1 (30.4 s) | Nemotron 3 (30.4 s) | Nemotron 3 (0.32 s) |
|---|---|---|---|
| DIHARD III full | 19.09 | 12.73 | 13.55 |
| DIHARD III 5–9 spk | 40.21 | 27.58 | 29.49 |
| CALLHOME-Part2 full | 10.32 | 9.10 | 11.32 |
| NOTSOFAR1 SC full | 30.49 | 11.00 | 14.53 |
| AMI Test SDM | 21.42 | 11.14 | 12.95 |
The biggest improvement is on recordings with five or more speakers. The card doesn’t state the old model’s speaker limit, but the 4spk in its name suggests it was built for four. The new model is also faster. On an RTX PRO 5000 in BF16, offline RTFx at batch size 32 is 12,196 vs 3,204 for v2.1 in eager mode, and 15,113 vs 2,619 with torch.compile.
Install
For NeMo, the card asks for Python 3.12 or later, Cython, and a recent PyTorch:
apt-get update && apt-get install -y libsndfile1 ffmpeguv pip install Cython packaginguv pip install 'nemo-toolkit[asr]'For the Transformers path, install Transformers from source:
pip install git+https://github.com/huggingface/transformersIf you only need a command-line tool, NeMo-Speech.cpp is a lightweight native C++ runtime. Follow its installation steps first, then run:
nemo-speech diarize meeting.wavnemo-speech transcribe meeting.wav --diarize --jsonRun it
This is the quick-start from the card, using NeMo. Its values are close to the card’s “Very high latency (offline)” row (30.4 s), but it doesn’t set spkcache_len, which is 264 in that row. The batch example below sets all five values.
from nemo.collections.asr.models import SortformerEncLabelModel
diar_model = SortformerEncLabelModel.from_pretrained("nvidia/Nemotron-3-Diarization")diar_model.eval()
diar_model.sortformer_modules.chunk_len = 340diar_model.sortformer_modules.chunk_right_context = 40diar_model.sortformer_modules.fifo_len = 40diar_model.sortformer_modules.spkcache_update_period = 300diar_model._check_streaming_parameters()
predicted_segments = diar_model.diarize(audio=["/path/to/your/audio.wav"], batch_size=1)
for segment in predicted_segments[0]: print(segment)Each segment comes back as begin_seconds, end_seconds, speaker_index.
Batch-diarize a folder of recordings
When you have a backlog of calls or meetings, diarize() accepts a line-delimited JSON manifest. Each line gives an offset and duration, so you can also process just one part of a long file. The example below builds a manifest with each file’s real length and runs every file through the offline configuration. The parameter values are copied from the card’s “Very high latency (offline)” row.
import jsonimport wavefrom pathlib import Path
from nemo.collections.asr.models import SortformerEncLabelModel
AUDIO_DIR = Path("/path/to/recordings")MANIFEST = Path("manifest.json")
def wav_duration(path: Path) -> float: """Length of a PCM WAV file in seconds (Python standard library).""" with wave.open(str(path), "rb") as w: return w.getnframes() / w.getframerate()
# Build a manifest: one JSON object per line, with a numeric duration as in the card.files = sorted(AUDIO_DIR.glob("*.wav"))with MANIFEST.open("w") as f: for wav in files: entry = {"audio_filepath": str(wav), "offset": 0, "duration": wav_duration(wav)} f.write(json.dumps(entry) + "\n")
diar_model = SortformerEncLabelModel.from_pretrained("nvidia/Nemotron-3-Diarization")diar_model.eval()
# Offline configuration (30.4 s input buffer), all values in 80 ms framesdiar_model.sortformer_modules.spkcache_len = 264diar_model.sortformer_modules.fifo_len = 40diar_model.sortformer_modules.chunk_len = 340diar_model.sortformer_modules.chunk_right_context = 40diar_model.sortformer_modules.spkcache_update_period = 300diar_model._check_streaming_parameters()
predicted_segments = diar_model.diarize(audio=str(MANIFEST), batch_size=8)
# Assumes results come back in manifest order (the card doesn't state this).for wav, segments in zip(files, predicted_segments): print(f"== {wav.name}") for segment in segments: print(segment)Python’s wave module only reads PCM WAV files. For other formats, get the duration some other way. The final loop pairs each result with its file by position. That assumes diarize() returns results in manifest order, which seems likely, but the card doesn’t say so. Larger batches are where most of the throughput comes from: at the offline setting the card reports RTFx of 1,340 at batch size 1 and 12,196 at batch size 32, both in eager mode.
Live speaker labels for a streaming pipeline
To use speaker labels in live captions or a voice agent, use the Transformers streaming API. You pick a streaming_mode, send the audio one chunk at a time, and pass speaker_cache from each forward call into the next one. This example follows the card. It simulates a stream from an audio file, so you would replace the slicing with your microphone or WebRTC buffer.
import torchfrom transformers import AutoModelForAudioFrameClassification, AutoProcessorfrom transformers.audio_utils import load_audio
model_id = "nvidia/Nemotron-3-Diarization"processor = AutoProcessor.from_pretrained(model_id)model = AutoModelForAudioFrameClassification.from_pretrained(model_id, device_map="auto")processor.set_streaming_mode("very_low_latency") # or "low_latency" (default), "ultra_low_latency"print(f"Streaming latency: {processor.streaming_latency_ms} ms")
sampling_rate = processor.feature_extractor.sampling_rateaudio = load_audio( "https://huggingface.co/datasets/hf-internal-testing/dummy-audio-samples/resolve/main/diarization_example.mp3", sampling_rate=sampling_rate,)
def inputs_generator(): yield processor( audio[: processor.num_samples_first_audio_chunk], sampling_rate=sampling_rate, is_streaming=True, is_first_audio_chunk=True, ) mel_frame_idx = processor.num_mel_frames_per_step start_idx = processor.audio_chunk_start(mel_frame_idx) while (end_idx := start_idx + processor.num_samples_per_audio_chunk) <= audio.shape[0]: yield processor( audio[start_idx:end_idx], sampling_rate=sampling_rate, is_streaming=True, is_first_audio_chunk=False, ) mel_frame_idx += processor.num_mel_frames_per_step start_idx = processor.audio_chunk_start(mel_frame_idx) yield processor( audio[start_idx:], sampling_rate=sampling_rate, is_streaming=True, is_first_audio_chunk=False, is_last_audio_chunk=True, )
speaker_cache, logits = None, []with torch.inference_mode(): for inputs in inputs_generator(): inputs = inputs.to(model.device, dtype=model.dtype) outputs = model(**inputs, speaker_cache=speaker_cache) logits.append(outputs.logits) # the chunk's frames, without its look-ahead speaker_cache = outputs.speaker_cache
logits = torch.cat(logits, dim=1) # (1, num_frames, 8), 10 ms framesfor seg in processor.extract_speaker_dict(logits)[0]: print(f"speaker_{seg['Speaker']}: {seg['Start']:.2f}s - {seg['End']:.2f}s")The card only shows how to turn the collected logits into segments with extract_speaker_dict. It doesn’t say whether the per-chunk logits are probabilities or raw scores, so if you want live per-chunk labels, check the Transformers documentation before you threshold them yourself. To pair diarization with streaming ASR, see the card’s ASR Integration Guide.
Talk-time analytics for recorded meetings
A common follow-up task is reporting how long each participant spoke. The offline Transformers path returns dicts with Start, End and Speaker keys, so the sums are simple:
from collections import defaultdict
import torchfrom transformers import AutoModelForAudioFrameClassification, AutoProcessorfrom transformers.audio_utils import load_audio
model_id = "nvidia/Nemotron-3-Diarization"processor = AutoProcessor.from_pretrained(model_id)model = AutoModelForAudioFrameClassification.from_pretrained(model_id, device_map="auto")
sampling_rate = processor.feature_extractor.sampling_rateaudio = load_audio( "https://huggingface.co/datasets/hf-internal-testing/dummy-audio-samples/resolve/main/diarization_example.mp3", sampling_rate=sampling_rate,)inputs = processor(audio, sampling_rate=sampling_rate).to(model.device, dtype=model.dtype)
with torch.inference_mode(): logits = model(**inputs).logits
segments = processor.extract_speaker_dict(logits, inputs.attention_mask)[0]
talk_time = defaultdict(float)turns = defaultdict(int)for seg in segments: talk_time[seg["Speaker"]] += seg["End"] - seg["Start"] turns[seg["Speaker"]] += 1
total = sum(talk_time.values())print(f"Detected {len(talk_time)} speakers")for spk in sorted(talk_time): share = 100 * talk_time[spk] / total if total else 0 print(f"speaker_{spk}: {talk_time[spk]:.1f}s ({share:.0f}%), {turns[spk]} turns")Overlapping speech counts toward every speaker who is talking at the time, so the percentages are shares of total speaker time, not of the recording’s length.
Gotchas
- Audio must be 16 kHz mono. If you pass numpy arrays to NeMo’s
diarize(), you must setsample_rateas an integer. The default is 16000. - Streaming parameters are measured in 80 ms frames, not seconds. Input buffer latency is
(CHUNK_LEN + RIGHT_CONTEXT) × 80 msand does not include compute time. Call_check_streaming_parameters()after you change them. - Lower latency costs accuracy in both DER and speaker counting. On NOTSOFAR1 SC, DER rises from 11.00 offline to 14.53 at 0.32 s, and speaker counting accuracy falls from 78.12 to 55.00. Throughput falls too: at 0.32 s, batch size 1, eager mode, RTFx is 12.5. 80 ms latency is possible, but the card recommends 0.32 s as the lowest setting.
- Speaker labels are generic. Channels are ordered by when each speaker first appears. The model does not identify people, and it handles at most eight speakers.
- It is not better everywhere. On AMI the card shows lower speaker counting accuracy than v2.1 (87.50 vs 93.75 offline). On two-speaker CALLHOME, DER is slightly worse than v2.1 at both 30.4 s (5.98 vs 5.68) and 1.04 s (6.98 vs 6.83).
- Reproducing DER depends on the reference labels. The AMI, AliMeeting and NOTSOFAR1 scores use forced-alignment reference RTTMs. Results scored against other labels are not comparable.
- Platform. The card lists Linux on NVIDIA Ampere, Ada, Hopper and Blackwell GPUs as supported, and lists no support for other platforms. It says nothing about which platforms NeMo-Speech.cpp runs on. Loading with
from_pretrainedmay require a Hugging Face token. The Transformers integration requires installing Transformers from source. - License. The model is released under OpenMDW-1.1, and the card says it is ready for commercial use. Read the license text before you ship it.
When to pick it
Pick it if you need diarization for meetings, calls or podcasts with more than four speakers, or if you want the same model for batch jobs and live streams. Its largest absolute DER gains over NVIDIA’s previous streaming Sortformer are on crowded meetings (offline NOTSOFAR1 SC 30.49 → 11.00, NOTSOFAR1 MHM 21.77 → 6.77). It is also faster at every setting the card reports. The speedup ranges from about 1.5× (offline, batch size 1, eager) to about 6× (1.04 s latency, batch size 32, compiled).
Look elsewhere if your recordings regularly have more than eight speakers, if you need to recognize specific named people instead of getting generic labels, or if you have to run on CPU, macOS or non-NVIDIA hardware, since the card lists no support for those. For mostly two-person phone calls, the card shows 2-speaker CALLHOME DER slightly worse than v2.1 at 30.4 s and 1.04 s, though full-set speaker counting improves.
Related
Run Audio8 ASR Infinite for 24/7 Chinese and English Transcription
A 3B-decoder streaming ASR model with 240–560 ms delay. With the authors' adapted vLLM build, a rolling KV cache lets it transcribe audio of any length at constant memory.
Edge0/Audio8-ASR-Infinite
Run Breeze TTS 2 for Voice Cloning, Design and Streaming
Breeze TTS 2 is an open-weight English and Chinese TTS model with voice cloning, text-described voices and a streaming API. Weights are non-commercial.
BreezeBlue/Breeze-TTS-2
Run Confucius4-R2T2 for Append-Only Real-Time Speech Recognition
NetEase Youdao's streaming ASR model built on Qwen3-ASR: transcript text never gets rewritten, with 80 ms to 2 s chunks. Setup, code and caveats.
netease-youdao/Confucius4-R2T2