SAVRN
Search Contact SAVRN

Open-weight model · Speech recognition

Voxtral-Mini-4B-Realtime-2602

by Mistral AI_ mistralai/Voxtral-Mini-4B-Realtime-2602

Voxtral Mini 4B Realtime 2602 is a multilingual, realtime speech-transcription model and among the first open-source solutions to achieve accuracy comparable to offline systems with a delay of = 3600 / 0.8 = 45000.

Parameters4.4B
Context131,072
Weights17.7 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads2M

Runs On

What it takes to serve Voxtral-Mini-4B-Realtime-2602 (4.4B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 8.9 GB 10.6 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 4.4 GB 5.3 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 2.2 GB 2.7 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

SAVRN's Notes on Voxtral-Mini-4B-Realtime-2602

Live transcription is the job here: 4.4B parameters from Mistral AI that take audio in and stream text out, with a 131,072-token context the publisher says covers about three hours under vLLM's defaults. At 16-bit it needs 10.6 GB of memory, 8.9 GB of that weights; 5.3 GB at 8-bit, 2.7 GB at 4-bit. The cheapest setup we price, one MI300X with 192 GB at $1.85 per hour on demand, sits nearly empty under one instance, so plan for many streams per accelerator.

Apache 2.0 asks only that you keep the notices and state significant changes, patent grant included, so commercial use is simple. Two checks before committing: it is derived from Ministral-3-3B-Base-2512, and the configuration lists an 8,192-token sliding window beside the long context, so test with the recording lengths you actually handle. No Index host prices it per token yet, so plan on the hourly card cost.

Model Card

By Mistral AI_, published under apache-2.0, revision 2769294da956.

Voxtral Mini 4B Realtime 2602 is a multilingual, realtime speech-transcription model and among the first open-source solutions to achieve accuracy comparable to offline systems with a delay of <500ms. It supports 13 languages and outperforms existing open-source baselines across a range of tasks, making it ideal for applications like voice assistants and live subtitling.

Built with a natively streaming architecture and a custom causal audio encoder - it allows configurable transcription delays (240ms to 2.4s), enabling users to balance latency and accuracy based on their needs. At a 480ms delay, it matches the performance of leading offline open-source transcription models, as well as realtime APIs.

As a 4B-parameter model, is optimized for on-device deployment, requiring minimal hardware resources. It runs in realtime with on devices minimal hardware with throughput exceeding 12.5 tokens/second.

This model is released in BF16 under the Apache-2 license, ensuring flexibility for both research and commercial use.

Read the full model card (1,407 words)

Configuration

Architecture
VoxtralRealtimeForConditionalGeneration
Context length (tokens)
131,072
Layers
26
Hidden size
3,072
Feed-forward size
9,216
Attention heads
32
Key/value heads
8
Head dimension
128
Vocabulary size
131,072
Sliding window (tokens)
8,192
Model type
voxtral_realtime

Identity and Version

Repository
mistralai/Voxtral-Mini-4B-Realtime-2602
Publisher
Mistral AI_
Task
Speech recognition
Modality
Audio
Library
vllm
Parameters
4.4B parameters
Languages
en, fr, es, de, ru, zh, ja, it
Revision
2769294da9567371363522aac9bbcfdd19447add
First published
2026-01-21
Last updated
2026-03-11

Files and Weights

9 files, 17.7 GB in total. The weights are 2 files totalling 17.7 GB in safetensors.

Weights2 files · 17.7 GB
Configuration5 files · 14.9 MB
Documentation1 file · 14.9 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
consolidated.safetensorsWeights8.9 GB 263f178fe752
model.safetensorsWeights8.9 GB e745e4902df6
config.jsonConfiguration1.6 KB
generation_config.jsonConfiguration191 B
params.jsonConfiguration1.3 KB
processor_config.jsonConfiguration384 B
tekken.jsonConfiguration14.9 MB 8434af1d39eb
README.mdDocumentation14.9 KB
.gitattributesRepository1.6 KB

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
17.7 GB
Download from Mistral AI_

Released by Mistral AI_ through its official repository on Hugging Face. Read the license.

Built From

  • Derived from mistralai/Ministral-3-3B-Base-2512
  • Described by arXiv:2602.11298

Evaluations

Each result is shown as reported, with the conditions its reporter stated. None is a SAVRN measurement. A comparison lines two results up only when their configuration, unit and setup are all stated and identical.

BenchmarkConditionsResultReported byRevisionDate
hf-audio/open-asr-leaderboard Task ami_werMetric ami_werComparison conditions not established 17.07 open-asr-leaderboard
Reported by a third party
Evaluated revision not stated 2026-01-21
hf-audio/open-asr-leaderboard Task earnings22_werMetric earnings22_werComparison conditions not established 11.84 open-asr-leaderboard
Reported by a third party
Evaluated revision not stated 2026-01-21
hf-audio/open-asr-leaderboard Task gigaspeech_werMetric gigaspeech_werComparison conditions not established 10.38 open-asr-leaderboard
Reported by a third party
Evaluated revision not stated 2026-01-21
hf-audio/open-asr-leaderboard Task librispeech_clean_werMetric librispeech_clean_werComparison conditions not established 2.08 open-asr-leaderboard
Reported by a third party
Evaluated revision not stated 2026-01-21
hf-audio/open-asr-leaderboard Task librispeech_other_werMetric librispeech_other_werComparison conditions not established 5.52 open-asr-leaderboard
Reported by a third party
Evaluated revision not stated 2026-01-21
hf-audio/open-asr-leaderboard Task mean_werMetric mean_werComparison conditions not established 7.68 open-asr-leaderboard
Reported by a third party
Evaluated revision not stated 2026-01-21
hf-audio/open-asr-leaderboard Task rtfxMetric rtfxComparison conditions not established 93.32 open-asr-leaderboard
Reported by a third party
Evaluated revision not stated 2026-01-21
hf-audio/open-asr-leaderboard Task spgispeech_werMetric spgispeech_werComparison conditions not established 2.42 open-asr-leaderboard
Reported by a third party
Evaluated revision not stated 2026-01-21
hf-audio/open-asr-leaderboard Task tedlium_werMetric tedlium_werComparison conditions not established 3.79 open-asr-leaderboard
Reported by a third party
Evaluated revision not stated 2026-01-21
hf-audio/open-asr-leaderboard Task voxpopuli_werMetric voxpopuli_werComparison conditions not established 8.34 open-asr-leaderboard
Reported by a third party
Evaluated revision not stated 2026-01-21

Memory Requirements

PrecisionWeights in memory
As published17.7 GB
16-bit8.9 GB
8-bit4.4 GB
4-bit2.2 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Compare Voxtral-Mini-4B-Realtime-2602

Questions About Voxtral-Mini-4B-Realtime-2602

How much GPU memory does Voxtral-Mini-4B-Realtime-2602 need?

About 10.6 GB at 16-bit and 2.7 GB at 4-bit: the weights (4.4B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run Voxtral-Mini-4B-Realtime-2602 on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use Voxtral-Mini-4B-Realtime-2602 commercially?

Yes. Voxtral-Mini-4B-Realtime-2602 is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

What is Voxtral-Mini-4B-Realtime-2602's context length?

131,072 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Speech recognition

Qwen3-ASR-1.7B

Qwen

The Qwen3-ASR family includes Qwen3-ASR-1.7B and Qwen3-ASR-0.6B, which support language identification and ASR for 52 languages and dialects. Both leverage large-scale speech training data and the strong audio understanding capability of their foundation model, Qwen3-Omni. Experiments show that the 1.7B version achieves state-of-the-art performance among open-source ASR models and is competitive with the strongest proprietary commercial APIs. Here are the main features: Novel and strong forced alignment Solution: We introduce Qwen3-ForcedAligner-0.6B, which supports timestamp prediction for arbitrary units within up to 5 minutes of speech in 11 languages. Evaluations show its timestamp…

Open weights apache-2.0 2.3B parameters

Model · Speech recognition

seamless-m4t-v2-large

AI at Meta

SeamlessM4T is our foundational all-in-one Massively Multilingual and Multimodal Machine Translation model delivering high-quality translation for speech and text in nearly 100 languages. SeamlessM4T models support the tasks of: - Automatic speech recognition (ASR). - 101 languages for speech input. - 96 Languages for text input/output. - 35 languages for speech output. We are releasing SeamlessM4T v2, an updated version with our novel UnitY2 architecture. This new model improves over SeamlessM4T v1 in quality as well as inference speed in speech generation tasks. The v2 version of SeamlessM4T is a multitask adaptation of our novel UnitY2 architecture. Unity2 with its hierarchical…

Open weights cc-by-nc-4.0 2.3B parameters 4,096 tokens transformers

Model · Speech recognition

cohere-transcribe-03-2026-mlx-4bit

David Larrea

Quantized MLX weights for beshkenadze/cohere-transcribe-03-2026-mlx-fp16. - model.safetensors - config.json - tokenizer.model - tokenizerconfig.json - preprocessorconfig.json - specialtokensmap.json - keymap.json - conversionsummary.json This checkpoint has been re-validated against the current Swift and Python MLX runtimes. Verified semantic parity on an English fixture: - official CUDA reference path (transformers native Cohere ASR) Fastest and smallest, but introduces a lexical regression on the repo sample (Kaldi → Khaldi). - Generated from the Swift-compatible fp16 checkpoint beshkenadze/cohere-transcribe-03-2026-mlx-fp16. - This repository contains inference artifacts only. Refer to…

Open weights apache-2.0 2.1B parameters mlx

Model · Speech recognition

cohere-transcribe-03-2026-mlx-8bit

David Larrea

Quantized MLX weights for beshkenadze/cohere-transcribe-03-2026-mlx-fp16. - model.safetensors - config.json - tokenizer.model - tokenizerconfig.json - preprocessorconfig.json - specialtokensmap.json - keymap.json - conversionsummary.json This checkpoint has been re-validated against the current Swift and Python MLX runtimes. Verified semantic parity on an English fixture: - official CUDA reference path (transformers native Cohere ASR) Matches fp16 on the repo sample while reducing memory substantially. - Generated from the Swift-compatible fp16 checkpoint beshkenadze/cohere-transcribe-03-2026-mlx-fp16. - This repository contains inference artifacts only. Refer to the upstream Cohere model…

Open weights apache-2.0 2.1B parameters mlx

Model · Speech recognition

whisper-large-v3

OpenAI

Whisper is a state-of-the-art model for automatic speech recognition (ASR) and speech translation, proposed in the paper Robust Speech Recognition via Large-Scale Weak Supervision by Alec Radford et al. from OpenAI. Trained on >5M hours of labeled data, Whisper demonstrates a strong ability to generalise to many datasets and domains in a zero-shot setting. Whisper large-v3 has the same architecture as the previous large and large-v2 models, except for the following minor differences: 1. The spectrogram input uses 128 Mel frequency bins instead of 80 The Whisper large-v3 model was trained on 1 million hours of weakly labeled audio and 4 million hours of pseudo-labeled audio collected using…

Open weights apache-2.0 1.5B parameters transformers

Model · Speech recognition

whisper-ja-1.5B

Efwkjn

For usage instructions follow openai/whisper-large-v3. Large-v3 finetune trained as a baseline with smaller checkpoints in progress. Expecting worse long form and equal short form. Benchmarks. Has occasional repetition issue compared to previous models but achieves competitive/SOTA CER across all tested sets.

Open weights 1.5B parameters