A lightweight standalone audio feature extractor derived from Voxtral-Mini-3B-2507. This repository isolates the Whisper-based audio encoder and multi-modal projector from the original 3B language model. The extracted module is intended for offline audio preprocessing, dataset preparation, feature caching, and downstream multimodal training pipelines. The language-model decoder is not included. It does not contain the 3B LLaMA language-model decoder. For offline preprocessing, loading the complete multimodal language model is unnecessary when the only required output is the projected audio representation. This standalone checkpoint can therefore be used as a dedicated audio feature…