SAVRN
Search Contact SAVRN

Open-weight model · Voice activity detection

Nemotron-3-Diarization-preview

by NVIDIA nvidia/Nemotron-3-Diarization-preview

For the detailed information, see the Overview subcard.

Parameters
Context
Weights397.2 MB
Licenseother
AccessAccess requested at publisher
Monthly Downloads1.1k

Model Card

For the detailed information, see the Overview subcard.

Excerpt from the card by NVIDIA, licensed other.

Identity and Version

Repository
nvidia/Nemotron-3-Diarization-preview
Publisher
NVIDIA
Task
Voice activity detection
Modality
Other
Library
nemo
Parameters
Not stated by the source
Languages
Not stated by the source
Revision
bcd3d20491b7a864c24a0abf65d8ed6b157c5e81
First published
2026-08-24
Last updated
2026-09-18

Files and Weights

11 files, 399.7 MB in total. The weights are 1 file totalling 397.2 MB in nemo.

Weights1 file · 397.2 MB
Documentation8 files · 64.5 KB
Other1 file · 2.5 MB
Repository1 file · 1.7 KB
Every file
FileTypeSizeSHA-256
Nemotron-3-Diarization-preview.nemoWeights397.2 MB
ASR_INTEGRATION_GUIDE.mdDocumentation16.4 KB
README.mdDocumentation2.3 KB
bias.mdDocumentation296 B
diarization_evaluation.mdDocumentation6.0 KB
explainability.mdDocumentation2.4 KB
overview.mdDocumentation35.2 KB
privacy.mdDocumentation1.3 KB
safety.mdDocumentation569 B
nemotron3_tts_8_open_voices_v18.mp4Other2.5 MB
.gitattributesRepository1.7 KB

License and Download

License
other
Access
Access requested at publisher
Download size
397.2 MB
Request access from NVIDIA

NVIDIA grants access through its official repository on Hugging Face.

Memory Requirements

PrecisionWeights in memory
As published397.2 MB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About Nemotron-3-Diarization-preview

What license is Nemotron-3-Diarization-preview released under?

other, as its publisher declares it. Read the license text before commercial use.

Similar Models

Model · Voice activity detection

segmentation-3.0

Pyannote

Using this open-source model in production? Consider switching to pyannoteAI for better and faster options. This model ingests 10 seconds of mono audio sampled at 16kHz and outputs speaker diarization as a (numframes, numclasses) matrix where the 7 classes are non-speech, speaker #1, speaker #2, speaker #3, speakers #1 and #2, speakers #1 and #3, and speakers #2 and #3. The various concepts behind this model are described in details in this paper. It has been trained by Séverin Baroudi with pyannote.audio 3.0.0 using the combination of the training sets of AISHELL, AliMeeting, AMI, AVA-AVD, DIHARD, Ego4D, MSDWild, REPERE, and VoxConverse. This companion repository by Alexis Plaquet also…

Access requested at publisher mit pyannote-audio

Model · Voice activity detection

segmentation

Pyannote

Using this open-source model in production? Consider switching to pyannoteAI for better and faster options. In order to reproduce the results of the paper "End-to-end speaker segmentation for overlap-aware resegmentation ", use pyannote/segmentation@Interspeech2021 with the following hyper-parameters: Expected outputs (and VBx baseline) are also provided in the /reproducibleresearch sub-directories.

Access requested at publisher mit pyannote-audio