SAVRN Model Hub · Comparisons
nemotron-3.5-asr-streaming-0.6b vs whisper-large-v3-turbo
Nemotron-3.5-asr-streaming-0.6b has 638M parameters and whisper-large-v3-turbo has 809M parameters; nemotron-3.5-asr-streaming-0.6b is released under other and whisper-large-v3-turbo under MIT License; at 16-bit, nemotron-3.5-asr-streaming-0.6b needs about 1.5 GB (1x MI300X from $1.85 an hour) and whisper-large-v3-turbo about 1.9 GB (1x MI300X from $1.85 an hour).
| Field | nemotron-3.5-asr-streaming-0.6b nvidia/nemotron-3.5-asr-streaming-0.6b | whisper-large-v3-turbo openai/whisper-large-v3-turbo |
|---|---|---|
| Publisher | NVIDIA | OpenAI |
| Task | Speech recognition | Speech recognition |
| Modality | Audio | Audio |
| Parameters, as reported | 638M parameters | 809M parameters |
| Architecture | Nemotron3_5AsrForRNNT | WhisperForConditionalGeneration |
| Library | nemo | transformers |
| Context length | Not stated | Not stated |
| Repository size | 5.7 GB | 1.6 GB |
| Artifact formats | safetensors, gguf, pytorch | safetensors |
| License | other | mit |
| Access | Open weights, no gate | Open weights, no gate |
| Memory at 16-bit (weights and margin) | 1.5 GB | 1.9 GB |
| Cheapest GPUs at 16-bit, per hour | 1x MI300X, $1.85 | 1x MI300X, $1.85 |
| Memory at 4-bit (weights and margin) | 0.4 GB | 0.5 GB |
| Cheapest GPUs at 4-bit, per hour | 1x MI300X, $1.85 | 1x MI300X, $1.85 |
| Revision viewed | ea30d66debe3 | 41f01f3fe87f |
| Downloads reported by the hub | 748.8k | 6.8M |
| Last observed | 2026-09-18 | 2026-09-18 |
An evaluation row appears only where at least two of these models report the same benchmark with the same stated configuration, metric, unit and setup. Different evaluators stay named in each cell. Values are shown as reported: no unit conversion, no ranking.
Other Reported Results
These results are listed for each model on its own, because the conditions needed to compare them are not stated or do not match. Two results that leave a condition blank are not assumed to share it.
nemotron-3.5-asr-streaming-0.6b
| Benchmark | Conditions | Result | Reported by | Revision | Date |
|---|---|---|---|---|---|
| ARTPARK-IISc/Vaani-Benchmark-V1.0 | Task Hindi_WERMetric Hindi_WERComparison conditions not established | 20.2 | Not named Reported by a third party |
Evaluated revision not stated | 2026-07-30 |
| FLEURS (English) | Configuration en_usTask Automatic Speech RecognitionMetric WER (1.12s frame size, LangID)Comparison conditions not established | 7.91 | nvidia Publisher reported |
Evaluated revision not stated | — |
| FLEURS (French) | Configuration fr_frTask Automatic Speech RecognitionMetric WER (1.12s frame size, LangID)Comparison conditions not established | 9.03 | nvidia Publisher reported |
Evaluated revision not stated | — |
| FLEURS (German) | Configuration de_deTask Automatic Speech RecognitionMetric WER (1.12s frame size, LangID)Comparison conditions not established | 8.31 | nvidia Publisher reported |
Evaluated revision not stated | — |
| FLEURS (Hindi) | Configuration hi_inTask Automatic Speech RecognitionMetric WER (1.12s frame size, LangID)Comparison conditions not established | 6.81 | nvidia Publisher reported |
Evaluated revision not stated | — |
| FLEURS (Italian) | Configuration it_itTask Automatic Speech RecognitionMetric WER (1.12s frame size, LangID)Comparison conditions not established | 4.25 | nvidia Publisher reported |
Evaluated revision not stated | — |
| FLEURS (Korean) | Configuration ko_krTask Automatic Speech RecognitionMetric WER (1.12s frame size, LangID)Comparison conditions not established | 7.12 | nvidia Publisher reported |
Evaluated revision not stated | — |
| FLEURS (Portuguese) | Configuration pt_brTask Automatic Speech RecognitionMetric WER (1.12s frame size, LangID)Comparison conditions not established | 5.48 | nvidia Publisher reported |
Evaluated revision not stated | — |
| FLEURS (Spanish) | Configuration es_419Task Automatic Speech RecognitionMetric WER (1.12s frame size, LangID)Comparison conditions not established | 4.11 | nvidia Publisher reported |
Evaluated revision not stated | — |
whisper-large-v3-turbo
| Benchmark | Conditions | Result | Reported by | Revision | Date |
|---|---|---|---|---|---|
| hf-audio/open-asr-leaderboard | Task ami_werMetric ami_werComparison conditions not established | 16.13 | open-asr-leaderboard Reported by a third party |
Evaluated revision not stated | 2024-10-01 |
| hf-audio/open-asr-leaderboard | Task earnings22_werMetric earnings22_werComparison conditions not established | 11.63 | open-asr-leaderboard Reported by a third party |
Evaluated revision not stated | 2024-10-01 |
| hf-audio/open-asr-leaderboard | Task gigaspeech_werMetric gigaspeech_werComparison conditions not established | 10.14 | open-asr-leaderboard Reported by a third party |
Evaluated revision not stated | 2024-10-01 |
| hf-audio/open-asr-leaderboard | Task librispeech_clean_werMetric librispeech_clean_werComparison conditions not established | 2.1 | open-asr-leaderboard Reported by a third party |
Evaluated revision not stated | 2024-10-01 |
| hf-audio/open-asr-leaderboard | Task librispeech_other_werMetric librispeech_other_werComparison conditions not established | 4.24 | open-asr-leaderboard Reported by a third party |
Evaluated revision not stated | 2024-10-01 |
| hf-audio/open-asr-leaderboard | Task mean_werMetric mean_werComparison conditions not established | 7.83 | open-asr-leaderboard Reported by a third party |
Evaluated revision not stated | 2024-10-01 |
| hf-audio/open-asr-leaderboard | Task rtfxMetric rtfxComparison conditions not established | 200.19 | open-asr-leaderboard Reported by a third party |
Evaluated revision not stated | 2024-10-01 |
| hf-audio/open-asr-leaderboard | Task spgispeech_werMetric spgispeech_werComparison conditions not established | 2.97 | open-asr-leaderboard Reported by a third party |
Evaluated revision not stated | 2024-10-01 |
| hf-audio/open-asr-leaderboard | Task tedlium_werMetric tedlium_werComparison conditions not established | 3.57 | open-asr-leaderboard Reported by a third party |
Evaluated revision not stated | 2024-10-01 |
| hf-audio/open-asr-leaderboard | Task voxpopuli_werMetric voxpopuli_werComparison conditions not established | 11.87 | open-asr-leaderboard Reported by a third party |
Evaluated revision not stated | 2024-10-01 |
SAVRN's Notes on nemotron-3.5-asr-streaming-0.6b
Live transcription is the job: audio in, punctuated and capitalized text out, with the chunk size set from 80 to 1,120 milliseconds. At 638 million parameters the 16-bit weights take 1.3 GB and need 1.5 GB to run, so the cheapest slot we list, one MI300X with 192 GB at $1.85 an hour, sits less than one percent occupied. That hour only pays if you stack streams on the card, and the 0.8 GB 8-bit and 0.4 GB 4-bit builds leave room for more.
The license reads other with no summary, so someone has to go through the publisher's terms before a commercial rollout. No context length is listed, so plan around chunk size. The training data names Common Voice 8.0, VoxPopuli, Europarl, FLEURS, Multilingual LibriSpeech and NVIDIA's Granary. And the weights changed on September 10, 2026, after the May 15 release, so pin the revision you validated.
SAVRN's Notes on whisper-large-v3-turbo
The lineage explains the size. This speech recognition build descends from openai/whisper-large-v3, pruned and fine-tuned with the decoding layers cut from 32 to 4, and lands at 809M parameters needing 1.9 GB at 16-bit, 1.0 GB at 8-bit or 0.5 GB at 4-bit. Against the cheapest listed rental, one 192 GB MI300X at $1.85 an hour, the design question we ask is how many audio streams to stack on one card, not whether it fits.
MIT asks for almost nothing: keep the copyright and permission notices, and commercial use, modification and redistribution are all yours. The reported evaluations are third-party figures and uneven by source, a mean word error rate of 7.83 with 2.1 on LibriSpeech clean and 16.13 on AMI, so pick the row that sounds like your audio. No context length is listed, so the audio window is a question for the publisher's documentation.
Questions
Which is larger, nemotron-3.5-asr-streaming-0.6b or whisper-large-v3-turbo?
whisper-large-v3-turbo (809M parameters) is larger than nemotron-3.5-asr-streaming-0.6b (638M parameters), by the parameter counts their publishers report.
Which is cheaper to run, nemotron-3.5-asr-streaming-0.6b or whisper-large-v3-turbo?
At 4-bit, nemotron-3.5-asr-streaming-0.6b fits on 1x MI300X from $1.85 an hour and whisper-large-v3-turbo on 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.
Can I use whisper-large-v3-turbo commercially?
Yes. whisper-large-v3-turbo is released under MIT License. The MIT License is a short permissive license. It permits commercial use, modification and redistribution, provided the copyright notice and permission notice are included.