Voxtral Mini Standalone Audio Feature Extractor
A lightweight standalone audio feature extractor derived from Voxtral-Mini-3B-2507.
This repository isolates the Whisper-based audio encoder and multi-modal projector from the original 3B language model. The extracted module is intended for offline audio preprocessing, dataset preparation, feature caching, and downstream multimodal training pipelines.
The language-model decoder is not included.
Voxtral Audio Tower Architecture
Raw audio
│
▼
HF Audio Feature Extractor
│
▼
Log-Mel features
│
▼
Voxtral Whisper-encoder
│
▼
[B, T, 1280]
│
▼
Feature packing ×4
│
▼
[B, T/4, 5120]
│
▼
Multi-Modal Projector
│
▼
[B, T/4, 3072]
The extracted audio branch contains:
- Voxtral / Whisper-based audio encoder
- Voxtral multi-modal projector
- Voxtral feature packing operation
It does not contain the 3B LLaMA language-model decoder.
Why?
For offline preprocessing, loading the complete multimodal language model is unnecessary when the only required output is the projected audio representation.
This standalone checkpoint can therefore be used as a dedicated audio feature extraction stage:
audio
↓
audio processor
↓
precomputed audio embeddings
↓
dataset cache
↓
LLM / adapter / multimodal training
This is particularly useful for large datasets where audio features can be computed once and reused across multiple training runs.
Output
For an input Mel tensor with shape:
[B, 128, T]
the extractor produces:
[B, T/4, 3072]
For example:
Input:
[1, 128, 1500]
Output:
[1, 375, 3072]
The current reference implementation uses bfloat16.
Quickstart
import torch
import soundfile as sf
import torchaudio.functional as F
from transformers import AutoModel, AutoFeatureExtractor
MODEL_ID = "vxltxr/voxtral-mini-audio-extractor"
device = "cuda" if torch.cuda.is_available() else "cpu"
# Standalone audio encoder + projector.
# No 3B LLM decoder is loaded.
model = AutoModel.from_pretrained(
MODEL_ID,
trust_remote_code=True,
dtype=torch.bfloat16,
).to(device).eval()
# Use the original Voxtral feature extractor for audio -> log-Mel preprocessing.
feature_extractor = AutoFeatureExtractor.from_pretrained(
"mistralai/Voxtral-Mini-3B-2507"
)
# Load audio.
audio_data, sampling_rate = sf.read(
"sample.wav",
dtype="float32",
)
# Convert stereo -> mono if necessary.
if audio_data.ndim > 1:
audio_data = audio_data.mean(axis=1)
# Resample to 16 kHz when necessary.
if sampling_rate != 16000:
waveform = torch.from_numpy(audio_data)
waveform = F.resample(
waveform,
orig_freq=sampling_rate,
new_freq=16000,
)
audio_data = waveform.numpy()
sampling_rate = 16000
# Audio -> log-Mel features.
inputs = feature_extractor(
audio_data,
sampling_rate=sampling_rate,
return_tensors="pt",
)
mel = inputs["input_features"].to(
device=device,
dtype=torch.bfloat16,
)
# Audio encoder -> packing -> multi-modal projector.
with torch.inference_mode():
audio_embeds = model.extract_features(mel)
print("Mel shape: ", mel.shape)
print("Embeddings shape:", audio_embeds.shape)
print("Embeddings dtype:", audio_embeds.dtype)
Offline Dataset Preprocessing
The intended use case is to precompute audio embeddings before training.
Conceptually:
def preprocess_batch(batch):
inputs = feature_extractor(
batch["audio"],
sampling_rate=16000,
return_tensors="pt",
)
mel = inputs["input_features"].to(
device="cuda",
dtype=torch.bfloat16,
)
with torch.inference_mode():
features = model.extract_features(mel)
return {
"audio_features": features.cpu().numpy(),
}
This allows the expensive audio encoder to run once during dataset preparation instead of during every training step.
Standalone Checkpoint
The extracted artifact contains the weights of:
audio_tower.*
multi_modal_projector.*
The extraction currently contains 489 tensors.
The standalone branch was validated against the corresponding reference audio graph with:
Output shape: [1, 375, 3072]
Max absolute diff: 0.0
Mean absolute diff: 0.0
Cosine similarity: 0.999999881
This verifies the extracted BF16 audio branch against the reference implementation for the tested input.
Relationship to Voxtral
This repository is derived from:
mistralai/Voxtral-Mini-3B-2507
The upstream Voxtral architecture combines an audio encoder and multi-modal projector with a language-model decoder. This repository extracts only the audio feature path for standalone preprocessing.
The Hugging Face Voxtral implementation describes get_audio_features() as the path that takes log-Mel audio features through the audio encoder and multi-modal projector to obtain audio embeddings.
Intended Use
Good fits include:
- offline audio feature extraction
- multimodal dataset preprocessing
- cached audio embeddings
- adapter / projector experiments
- multimodal LLM training pipelines
- large-scale dataset preparation
- debugging and analysis of the Voxtral audio branch
Not Intended For
This checkpoint is not a speech-to-text model.
It does not contain:
- the 3B LLaMA decoder
- text generation weights
- a tokenizer for generation
- the full Voxtral conditional-generation pipeline
Its output is an intermediate audio representation intended to be consumed by a downstream model.
Notes
The current checkpoint expects the Voxtral-compatible audio preprocessing pipeline to produce the appropriate log-Mel input_features.
For production dataset preprocessing, keep the audio preprocessing configuration aligned with the original Voxtral model.
License
Apache-2.0.
This repository contains extracted components derived from the upstream Voxtral model. Please also review the upstream model's license and terms before redistribution or deployment.