An equal-weight average of the epoch 6, 7, 8, 9 and 10 checkpoints of No training was done here, and inference costs exactly what one model costs. Like its ingredient, this model has no reported WER or CER and cannot have one. That run trains on every held-out hour the project has, including the SAPC2 dev split the rest of this family scores against. Averaging its epochs does not create a set to measure on. So this checkpoint rests on a bet rather than a measurement, and it is worth same soup of the same five epochs was worth 0.27 CER points — 6.06% against 6.33% for the best single epoch. That is the whole of the evidence. It is evidence from a different architecture (RNN-T, not TDT) on…
Open-weight model · Speech recognition
parakeet-tdt-0.6b-v3
by MLX Community mlx-community/parakeet-tdt-0.6b-v3
This model was converted to MLX format from nvidia/parakeet-tdt-0.6b-v3 using the conversion script. Please refer to original model card for more details on the model.
Runs On
What it takes to serve parakeet-tdt-0.6b-v3 (627M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.
| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
|---|---|---|---|---|---|
| 16-bit | 1.3 GB | 1.5 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 8-bit | 0.6 GB | 0.8 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 4-bit | 0.3 GB | 0.4 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.
SAVRN's Notes on parakeet-tdt-0.6b-v3
The job is speech recognition, and at 627M parameters the memory math is trivial: 1.5 GB at 16-bit, 0.8 GB at 8-bit, 0.4 GB at 4-bit. The cheapest Index listing that fits is one MI300X with 192 GB at $1.85 an hour on demand, so the hardware question is not which card but whether transcription needs its own card. The files are packaged for the mlx library in safetensors and mlx formats, so confirm your serving stack reads them.
CC BY 4.0 asks little: commercial use is allowed, you credit the creator and note any changes. Two checks before committing. This is a conversion, published by the MLX Community as derived from nvidia/parakeet-tdt-0.6b-v3, and the publisher points back to the original card for details, so verify behavior there. With no context length or host token price on file, budget the run in GPU hours, not tokens.
Model Card
By MLX Community, published under cc-by-4.0, revision ed2b7e8c15f9.
This model was converted to MLX format from nvidia/parakeet-tdt-0.6b-v3 using the conversion script. Please refer to original model card for more details on the model.
Read MLX Community's full model card
This model was converted to MLX format from nvidia/parakeet-tdt-0.6b-v3 using the conversion script. Please refer to original model card for more details on the model.
Use with mlx
parakeet-mlx
pip install -U parakeet-mlx
parakeet-mlx audio.wav --model mlx-community/parakeet-tdt-0.6b-v3
mlx-audio
pip install -U mlx-audio
python -m mlx_audio.stt.generate --model mlx-community/parakeet-tdt-0.6b-v3 --audio audio.wav --output somewhere
Identity and Version
- Repository
- mlx-community/parakeet-tdt-0.6b-v3
- Publisher
- MLX Community
- Task
- Speech recognition
- Modality
- Audio
- Library
- mlx
- Parameters
- 627M parameters
- Languages
- en, es, fr, de, bg, hr, cs, da
- Revision
- ed2b7e8c15f9aaa0b5772e2efb986255eaef7e15
- First published
- 2025-08-16
- Last updated
- 2025-08-16
Files and Weights
7 files, 2.5 GB in total. The weights are 1 file totalling 2.5 GB in safetensors.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| model.safetensors | Weights | 2.5 GB | 05e01c7f396c |
| config.json | Configuration | 244.1 KB | — |
| README.md | Documentation | 1.1 KB | — |
| .gitattributes | Repository | 1.5 KB | — |
| tokenizer.model | Tokenizer | 360.9 KB | eacec2b0a77f |
| tokenizer.vocab | Tokenizer | 101.0 KB | — |
| vocab.txt | Tokenizer | 46.8 KB | — |
License and Download
- License
- cc-by-4.0
- Access
- Open weights, no gate
- Download size
- 2.5 GB
Released by MLX Community through its official repository on Hugging Face. Read the license.
Built From
- Derived from nvidia/parakeet-tdt-0.6b-v3
Memory Requirements
| Precision | Weights in memory |
|---|---|
| As published | 2.5 GB |
| 16-bit | 1.3 GB |
| 8-bit | 0.6 GB |
| 4-bit | 0.3 GB |
Weights only, from the published parameter count; the key-value cache and runtime add to this.
Compare parakeet-tdt-0.6b-v3
Questions About parakeet-tdt-0.6b-v3
How much GPU memory does parakeet-tdt-0.6b-v3 need?
About 1.5 GB at 16-bit and 0.4 GB at 4-bit: the weights (627M parameters) plus a working margin. A long context needs more.
What is the cheapest GPU to run parakeet-tdt-0.6b-v3 on?
At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.
Can I use parakeet-tdt-0.6b-v3 commercially?
Yes. parakeet-tdt-0.6b-v3 is released under Creative Commons Attribution 4.0. CC BY 4.0 permits sharing and adapting the work, including commercially, provided the creator is credited and changes are indicated.
Similar Models
nvidia/parakeet-tdt-0.6b-v3 fine-tuned on all 1,047.7 hours this project holds: SAPC1 train and dev, SAPC2 train, the SAPC2 dev split the rest of this family scores against, 103.1 hours of synthetic dysarthric speech, 79.2 hours recovered by force-aligning and cutting recordings past the 45-second training cap, and 15.5 hours of AtaxiaUK and HeyJay!, which are outside the challenge corpora and make this an unconstrained-track model. This model has no reported WER or CER, and cannot have one. Every held-out hour is in its training data. That was the point: the hyperparameters were settled on the sibling runs that do hold out a dev split, and this run spends that split as training data…
This model was converted to MLX format from nvidia/parakeet-tdt-0.6b-v2 using the conversion script. Please refer to original model card for more details on the model.
/ Improve list spacing / / Badge alignment consistency / Nemotron 3.5 ASR is a multilingual, streaming Automatic Speech Recognition (ASR) model engineered to deliver high-quality multilingual transcription across both low-latency streaming and high-throughput batch workloads. Developed by NVIDIA, this 600M parameter model transcribes speech into text with native support for punctuation and capitalization, and offers runtime flexibility with configurable chunk sizes, including 80ms, 160ms, 320ms, 560ms, and 1120ms. By leveraging a state-of-the-art Cache-Aware FastConformer-RNNT architecture, the model eliminates redundant overlapping computations common in traditional "buffered" streaming.…
Whisper finetune for Japanese focused on general/anime domains. For usage instructions follow openai/whisper-large-v3-turbo. Due to vocab changes ctranslate2>=4.7.1 required for faster-whisper. For inference engines with hardcoded vocab, the token embedding can be padded. Finetuned from turbo with pruned vocab, indices can be found in mapping.txt. Trained decoder only for 2^20 steps, batch size 64. Using a 45000 hour corpus (largest source is 17000 of filtered reazonspeech-all) with custom mixing ratio and augmentation to maintain long form performance and timestamps. Benchmarks. Competitive/SOTA on test sets, slightly better than 1.5B on short form, worse on long form. Also trained for…
For usage instructions follow openai/whisper-large-v3-turbo. Note for faster-whisper vocab changes make model.ismultilingual and suppresstokens wrong. Please adjust the code as required if you want to use this with faster-whisper. Turbo finetune with japanese tokenizer. Full finetune trained 2^19 steps, batch size 64. Smaller vocab with ~1.6x bytes/token allows faster speed with 4 layers vs 2 layer distil (10% larger decoder). Benchmarks. Short form slightly behind v0.2 (trained less?) but long form much better. Also trained for lyrics but untested. Research supported with Cloud TPUs from Google's TPU Research Cloud (TRC)