SAVRN
Search Contact SAVRN

Open-weight model · Audio classification

mms-lid-256

by AI at Meta facebook/mms-lid-256

This checkpoint is a model fine-tuned for speech language identification (LID) and part of Facebook's Massive Multilingual Speech project.

Parameters966M
Context
Weights7.7 GB
Licensecc-by-nc-4.0
AccessOpen weights
Monthly Downloads38.3k

Runs On

What it takes to serve mms-lid-256 (966M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 1.9 GB 2.3 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 1.0 GB 1.2 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 0.5 GB 0.6 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

Model Card

This checkpoint is a model fine-tuned for speech language identification (LID) and part of Facebook's Massive Multilingual Speech project. This checkpoint is based on the Wav2Vec2 architecture and classifies raw audio input to a probability distribution over 256 output classes (each class representing a language). The checkpoint consists of 1 billion parameters and has been fine-tuned from facebook/mms-1b on 256 languages. This MMS checkpoint can be used with Transformers to identify the spoken language of an audio. It can recognize the following 256 languages. Let's look at a simple example. First, we install transformers and some other libraries Note: In order to use MMS you need to have…

Excerpt from the card by AI at Meta, licensed cc-by-nc-4.0.

Configuration

Architecture
Wav2Vec2ForSequenceClassification
Layers
48
Hidden size
1,280
Feed-forward size
5,120
Attention heads
16
Vocabulary size
154
Stored precision
float32
Model type
wav2vec2

Identity and Version

Repository
facebook/mms-lid-256
Publisher
AI at Meta
Task
Audio classification
Modality
Audio
Library
transformers
Parameters
966M parameters
Languages
ab, af, ak, am, ar, as, av, ay
Revision
edc73fd00996e671dfc59d16436a29b12b10588a
First published
2023-06-13
Last updated
2023-06-13

Files and Weights

7 files, 7.7 GB in total. The weights are 2 files totalling 7.7 GB in bin, safetensors.

Weights2 files · 7.7 GB
Configuration2 files · 6.8 KB
Documentation1 file · 7.9 KB
Other1 file · 1.5 KB
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
model.safetensorsWeights3.9 GB a946279c9911
pytorch_model.binWeights3.9 GB 258e7f82f96c
config.jsonConfiguration6.6 KB
preprocessor_config.jsonConfiguration212 B
README.mdDocumentation7.9 KB
langs.txtOther1.5 KB
.gitattributesRepository1.5 KB

License and Download

License
cc-by-nc-4.0
Access
Open weights, no gate
Download size
7.7 GB
Download from AI at Meta

Released by AI at Meta through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published7.7 GB
16-bit1.9 GB
8-bit1.0 GB
4-bit0.5 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About mms-lid-256

How much GPU memory does mms-lid-256 need?

About 2.3 GB at 16-bit and 0.6 GB at 4-bit: the weights (966M parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run mms-lid-256 on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use mms-lid-256 commercially?

Not without separate permission. mms-lid-256 is released under Creative Commons Attribution-NonCommercial 4.0. CC BY-NC 4.0 permits sharing and adapting with credit for non-commercial purposes only. Commercial use needs separate permission from the rights holder.

Similar Models

Model · Audio classification

mms-lid-126

AI at Meta

This checkpoint is a model fine-tuned for speech language identification (LID) and part of Facebook's Massive Multilingual Speech project. This checkpoint is based on the Wav2Vec2 architecture and classifies raw audio input to a probability distribution over 126 output classes (each class representing a language). The checkpoint consists of 1 billion parameters and has been fine-tuned from facebook/mms-1b on 126 languages. This MMS checkpoint can be used with Transformers to identify the spoken language of an audio. It can recognize the following 126 languages. Let's look at a simple example. First, we install transformers and some other libraries Note: In order to use MMS you need to have…

Open weights cc-by-nc-4.0 966M parameters transformers

Model · Audio classification

mms-lid-512

AI at Meta

This checkpoint is a model fine-tuned for speech language identification (LID) and part of Facebook's Massive Multilingual Speech project. This checkpoint is based on the Wav2Vec2 architecture and classifies raw audio input to a probability distribution over 512 output classes (each class representing a language). The checkpoint consists of 1 billion parameters and has been fine-tuned from facebook/mms-1b on 512 languages. This MMS checkpoint can be used with Transformers to identify the spoken language of an audio. It can recognize the following 512 languages. Let's look at a simple example. First, we install transformers and some other libraries Note: In order to use MMS you need to have…

Open weights cc-by-nc-4.0 966M parameters transformers

Model · Audio classification

mms-lid-1024

AI at Meta

This checkpoint is a model fine-tuned for speech language identification (LID) and part of Facebook's Massive Multilingual Speech project. This checkpoint is based on the Wav2Vec2 architecture and classifies raw audio input to a probability distribution over 1024 output classes (each class representing a language). The checkpoint consists of 1 billion parameters and has been fine-tuned from facebook/mms-1b on 1024 languages. This MMS checkpoint can be used with Transformers to identify the spoken language of an audio. It can recognize the following 1024 languages. Let's look at a simple example. First, we install transformers and some other libraries Note: In order to use MMS you need to…

Open weights cc-by-nc-4.0 967M parameters transformers

Model · Audio classification

mms-lid-4017

AI at Meta

This checkpoint is a model fine-tuned for speech language identification (LID) and part of Facebook's Massive Multilingual Speech project. This checkpoint is based on the Wav2Vec2 architecture and classifies raw audio input to a probability distribution over 4017 output classes (each class representing a language). The checkpoint consists of 1 billion parameters and has been fine-tuned from facebook/mms-1b on 4017 languages. This MMS checkpoint can be used with Transformers to identify the spoken language of an audio. It can recognize the following 4017 languages. Let's look at a simple example. First, we install transformers and some other libraries Note: In order to use MMS you need to…

Open weights cc-by-nc-4.0 970M parameters transformers

This project leverages the Whisper model to recognize emotions in speech. The goal is to classify audio recordings into different emotional categories, such as Happy, Sad, Surprised, and etc. The dataset used for training and evaluation is sourced from multiple datasets, including: The dataset contains recordings labeled with various emotions. Below is the distribution of the emotions in the dataset: This distribution reflects the balance of emotions in the dataset, with some emotions having more samples than others. Excluded the "calm" emotion during training due to its underrepresentation. The model used is the Whisper Large V3 model, fine-tuned for audio classification tasks: I map the…

Open weights apache-2.0 637M parameters transformers

Model · Audio classification

Qwen3-ForcedAligner-0.6B-4bit

Ivan

4-bit quantized version of Qwen/Qwen3-ForcedAligner-0.6B for Apple Silicon inference via MLX. Predicts word-level timestamps for audio+text pairs in a single non-autoregressive forward pass. Unlike ASR (autoregressive, token-by-token), the forced aligner runs the entire sequence in one forward pass through the decoder. The classify head predicts a timestamp class (0–4999) at each token position, which maps to time via classindex × 80ms. This model is designed for use with speech-swift: Text decoder (attention projections, MLP, embeddings) quantized to 4-bit using group quantization (groupsize=64). Audio encoder and classify head kept as float16 for accuracy.

Open weights apache-2.0 415M parameters mlx