SAVRN
Search Contact SAVRN

Open-weight model · Speech recognition

wav2vec2-base-960h

by AI at Meta facebook/wav2vec2-base-960h

The base model pretrained and fine-tuned on 960 hours of Librispeech on 16kHz sampled speech audio. When using the model make sure that your speech input is also sampled at 16Khz.

Parameters94M
Context
Weights1.1 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads1.5M

Runs On

What it takes to serve wav2vec2-base-960h (94M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 0.2 GB 0.2 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 0.1 GB 0.1 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 0.0 GB 0.1 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

SAVRN's Notes on wav2vec2-base-960h

If your audio arrives at 16 kHz, this AI at Meta checkpoint transcribes it with 94 million parameters and 0.2 GB of memory at 16-bit, 0.1 GB at 8-bit or 4-bit. Cheapest on the Index: one MI300X with 192 GB at $1.85 per hour, a card that could hold hundreds of copies, so budget by the audio pipeline rather than the accelerator. It has no token context length; it was trained on 960 hours of Librispeech, and the publisher reports a 3.4 word error rate on the clean test set and 8.6 on other.

Apache 2.0 allows commercial use, modification and redistribution with notices kept and significant changes stated, and carries an express patent grant, so it can ship inside a product. Check the third-party numbers first: the open-asr-leaderboard reports a mean word error rate of 29.4, with 45.56 on AMI, so test your own recordings before committing.

Model Card

By AI at Meta, published under apache-2.0, revision 22aad52d435e.

Facebook's Wav2Vec2

The base model pretrained and fine-tuned on 960 hours of Librispeech on 16kHz sampled speech audio. When using the model make sure that your speech input is also sampled at 16Khz.

Paper

Authors: Alexei Baevski, Henry Zhou, Abdelrahman Mohamed, Michael Auli

Abstract

Read the full model card (352 words)

Configuration

Architecture
Wav2Vec2ForCTC
Layers
12
Hidden size
768
Feed-forward size
3,072
Attention heads
12
Vocabulary size
32
Model type
wav2vec2

Identity and Version

Repository
facebook/wav2vec2-base-960h
Publisher
AI at Meta
Task
Speech recognition
Modality
Audio
Library
transformers
Parameters
94M parameters
Languages
en
Revision
22aad52d435eb6dbaf354bdad9b0da84ce7d6156
First published
2022-03-02
Last updated
2022-11-14

Files and Weights

11 files, 1.1 GB in total. The weights are 3 files totalling 1.1 GB in bin, h5, safetensors.

Weights3 files · 1.1 GB
Configuration4 files · 2.0 KB
Tokenizer2 files · 454 B
Documentation1 file · 4.4 KB
Repository1 file · 790 B
Every file
FileTypeSizeSHA-256
model.safetensorsWeights377.6 MB 8aa76ab2243c
pytorch_model.binWeights377.7 MB c34f9827b034
tf_model.h5Weights377.8 MB 412742825972
config.jsonConfiguration1.6 KB
feature_extractor_config.jsonConfiguration158 B
preprocessor_config.jsonConfiguration159 B
special_tokens_map.jsonConfiguration85 B
README.mdDocumentation4.4 KB
.gitattributesRepository790 B
tokenizer_config.jsonTokenizer163 B
vocab.jsonTokenizer291 B

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
1.1 GB
Download from AI at Meta

Released by AI at Meta through its official repository on Hugging Face. Read the license.

Built From

Evaluations

Each result is shown as reported, with the conditions its reporter stated. None is a SAVRN measurement. A comparison lines two results up only when their configuration, unit and setup are all stated and identical.

BenchmarkConditionsResultReported byRevisionDate
LibriSpeech (clean) Configuration cleanTask Automatic Speech RecognitionMetric Test WERComparison conditions not established 3.4 facebook
Publisher reported
Evaluated revision not stated
LibriSpeech (other) Configuration otherTask Automatic Speech RecognitionMetric Test WERComparison conditions not established 8.6 facebook
Publisher reported
Evaluated revision not stated
hf-audio/open-asr-leaderboard Task ami_werMetric ami_werComparison conditions not established 45.56 open-asr-leaderboard
Reported by a third party
Evaluated revision not stated 2022-03-02
hf-audio/open-asr-leaderboard Task earnings22_werMetric earnings22_werComparison conditions not established 48.47 open-asr-leaderboard
Reported by a third party
Evaluated revision not stated 2022-03-02
hf-audio/open-asr-leaderboard Task gigaspeech_werMetric gigaspeech_werComparison conditions not established 30.85 open-asr-leaderboard
Reported by a third party
Evaluated revision not stated 2022-03-02
hf-audio/open-asr-leaderboard Task librispeech_clean_werMetric librispeech_clean_werComparison conditions not established 12.53 open-asr-leaderboard
Reported by a third party
Evaluated revision not stated 2022-03-02
hf-audio/open-asr-leaderboard Task librispeech_other_werMetric librispeech_other_werComparison conditions not established 16.72 open-asr-leaderboard
Reported by a third party
Evaluated revision not stated 2022-03-02
hf-audio/open-asr-leaderboard Task mean_werMetric mean_werComparison conditions not established 29.4025 open-asr-leaderboard
Reported by a third party
Evaluated revision not stated 2022-03-02
hf-audio/open-asr-leaderboard Task rtfxMetric rtfxComparison conditions not established 686.003 open-asr-leaderboard
Reported by a third party
Evaluated revision not stated 2022-03-02
hf-audio/open-asr-leaderboard Task spgispeech_werMetric spgispeech_werComparison conditions not established 27.56 open-asr-leaderboard
Reported by a third party
Evaluated revision not stated 2022-03-02
hf-audio/open-asr-leaderboard Task tedlium_werMetric tedlium_werComparison conditions not established 21.05 open-asr-leaderboard
Reported by a third party
Evaluated revision not stated 2022-03-02
hf-audio/open-asr-leaderboard Task voxpopuli_werMetric voxpopuli_werComparison conditions not established 32.48 open-asr-leaderboard
Reported by a third party
Evaluated revision not stated 2022-03-02

Memory Requirements

PrecisionWeights in memory
As published1.1 GB
16-bit0.2 GB
8-bit0.1 GB
4-bit0.0 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Built on This Model

Questions About wav2vec2-base-960h

How much GPU memory does wav2vec2-base-960h need?

About 0.2 GB at 16-bit and 0.1 GB at 4-bit: the weights (94M parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run wav2vec2-base-960h on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use wav2vec2-base-960h commercially?

Yes. wav2vec2-base-960h is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Similar Models

Model · Speech recognition

whisper-base

OpenAI

Whisper is a pre-trained model for automatic speech recognition (ASR) and speech translation. Trained on 680k hours of labelled data, Whisper models demonstrate a strong ability to generalise to many datasets and domains without the need for fine-tuning. Whisper was proposed in the paper Robust Speech Recognition via Large-Scale Weak Supervision by Alec Radford et al from OpenAI. The original code repository can be found here. Disclaimer: Content for this model card has partly been written by the Hugging Face team, and parts of it were copied and pasted from the original model card. Whisper is a Transformer based encoder-decoder model, also referred to as a sequence-to-sequence model. It…

Open weights apache-2.0 73M parameters transformers

Model · Speech recognition

whisper-ja-51M

Efwkjn

Whisper finetune for Japanese focused on general/anime domains. For usage instructions follow openai/whisper-large-v3-turbo. Due to vocab changes ctranslate2>=4.7.1 required for faster-whisper. For inference engines with hardcoded vocab, the token embedding can be padded. Finetuned from base with pruned vocab and encoder conv adaption, indices can be found in mapping.txt. Trained decoder only for 2^20 steps, batch size 64. Using a 45000 hour corpus (largest source is 17000 of filtered reazonspeech-all) with custom mixing ratio and augmentation to maintain long form performance and timestamps. Benchmarks. Competitive for size on test sets, particually good on JSUT-book. Also trained for…

Open weights 51M parameters

Model · Speech recognition

whisper-tiny

OpenAI

Whisper is a pre-trained model for automatic speech recognition (ASR) and speech translation. Trained on 680k hours of labelled data, Whisper models demonstrate a strong ability to generalise to many datasets and domains without the need for fine-tuning. Whisper was proposed in the paper Robust Speech Recognition via Large-Scale Weak Supervision by Alec Radford et al from OpenAI. The original code repository can be found here. Disclaimer: Content for this model card has partly been written by the Hugging Face team, and parts of it were copied and pasted from the original model card. Whisper is a Transformer based encoder-decoder model, also referred to as a sequence-to-sequence model. It…

Open weights apache-2.0 38M parameters transformers

Model · Speech recognition

whisper-ja-22M

Efwkjn

Whisper finetune for Japanese focused on general/anime domains. For usage instructions follow openai/whisper-large-v3-turbo. Due to vocab changes ctranslate2>=4.7.1 required for faster-whisper. For inference engines with hardcoded vocab, the token embedding can be padded. Finetuned from tiny with pruned vocab and encoder conv adaption, indices can be found in mapping.txt. Trained decoder only for 2^19 steps, batch size 64. Using a 45000 hour corpus (largest source is 17000 of filtered reazonspeech-all) with custom mixing ratio and augmentation to maintain long form performance and timestamps. Benchmarks. CER roughly between OpenAI whisper-base/small, not great but also the smallest model…

Open weights 22M parameters

Model · Speech recognition

wav2vec2-large-xlsr-53-japanese

Jonatas Grosman

Fine-tuned facebook/wav2vec2-large-xlsr-53 on Japanese using the train and validation splits of Common Voice 6.1, CSS10 and JSUT. When using this model, make sure that your speech input is sampled at 16kHz. This model has been fine-tuned thanks to the GPU credits generously given by the OVHcloud:) The script used for training can be found here: https://github.com/jonatasgrosman/wav2vec2-sprint The model can be used directly (without a language model) as follows... Using the HuggingSound library: The model can be evaluated as follows on the Japanese test data of Common Voice. In the table below I report the Word Error Rate (WER) and the Character Error Rate (CER) of the model. I ran the…

Open weights apache-2.0 transformers

Model · Speech recognition

whisperkit-coreml

Argmax

WhisperKit is part of Argmax OSS, an On-device Speech AI SDK for Apple Silicon: https://github.com/argmaxinc/argmax-oss-swift Check out the WhisperKit paper and presentation from ICML 2025: https://icml.cc/virtual/2025/47854 For real-time transcription with speakers and custom vocabulary, check out Argmax Pro SDK: https://www.argmaxinc.com/blog/argmax-sdk-2

Open weights mit whisperkit