SAVRN
Search Contact SAVRN

Open-weight model · Audio classification

wav2vec2-lg-xlsr-en-speech-emotion-recognition

by Enrique Hernández Calabrés ehcalabres/wav2vec2-lg-xlsr-en-speech-emotion-recognition

The model is a fine-tuned version of jonatasgrosman/wav2vec2-large-xlsr-53-english for a Speech Emotion Recognition (SER) task. The dataset used to fine-tune the original pre-trained model is the RAVDESS dataset.

Parameters316M
Context
Weights2.5 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads25.6k

Runs On

What it takes to serve wav2vec2-lg-xlsr-en-speech-emotion-recognition (316M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 0.6 GB 0.8 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 0.3 GB 0.4 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 0.2 GB 0.2 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

Model Card

By Enrique Hernández Calabrés, published under apache-2.0, revision b520c9c46a71.

The model is a fine-tuned version of jonatasgrosman/wav2vec2-large-xlsr-53-english for a Speech Emotion Recognition (SER) task. The dataset used to fine-tune the original pre-trained model is the RAVDESS dataset. This dataset provides 1440 samples of recordings from actors performing on 8 different emotions in English, which are: It achieves the following results on the evaluation set: The following hyperparameters were used during training: - learningrate: 0.0001 - trainbatchsize: 4 - evalbatchsize: 4 - gradientaccumulationsteps: 2 - totaltrainbatchsize: 8 - lrschedulertype: linear - numepochs: 3 - mixedprecisiontraining: Native AMP Any doubt, contact me on Twitter. - Transformers 4.8.2…

Read Enrique Hernández Calabrés's full model card

Speech Emotion Recognition By Fine-Tuning Wav2Vec 2.0

The model is a fine-tuned version of jonatasgrosman/wav2vec2-large-xlsr-53-english for a Speech Emotion Recognition (SER) task.

The dataset used to fine-tune the original pre-trained model is the RAVDESS dataset. This dataset provides 1440 samples of recordings from actors performing on 8 different emotions in English, which are:

emotions = ['angry', 'calm', 'disgust', 'fearful', 'happy', 'neutral', 'sad', 'surprised']

It achieves the following results on the evaluation set: - Loss: 0.5023 - Accuracy: 0.8223

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training: - learning_rate: 0.0001 - train_batch_size: 4 - eval_batch_size: 4 - seed: 42 - gradient_accumulation_steps: 2 - total_train_batch_size: 8 - optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08 - lr_scheduler_type: linear - num_epochs: 3 - mixed_precision_training: Native AMP

Training results

Training Loss Epoch Step Validation Loss Accuracy
2.0752 0.21 30 2.0505 0.1359
2.0119 0.42 60 1.9340 0.2474
1.8073 0.63 90 1.5169 0.3902
1.5418 0.84 120 1.2373 0.5610
1.1432 1.05 150 1.1579 0.5610
0.9645 1.26 180 0.9610 0.6167
0.8811 1.47 210 0.8063 0.7178
0.8756 1.68 240 0.7379 0.7352
0.8208 1.89 270 0.6839 0.7596
0.7118 2.1 300 0.6664 0.7735
0.4261 2.31 330 0.6058 0.8014
0.4394 2.52 360 0.5754 0.8223
0.4581 2.72 390 0.4719 0.8467
0.3967 2.93 420 0.5023 0.8223

Citation

@misc {enrique_hernández_calabrés_2024,
    author       = { {Enrique Hernández Calabrés} },
    title        = { wav2vec2-lg-xlsr-en-speech-emotion-recognition (Revision 17cf17c) },
    year         = 2024,
    url          = { https://huggingface.co/ehcalabres/wav2vec2-lg-xlsr-en-speech-emotion-recognition },
    doi          = { 10.57967/hf/2045 },
    publisher    = { Hugging Face }
}

Contact

Any doubt, contact me on Twitter.

Framework versions

  • Transformers 4.8.2
  • Pytorch 1.9.0+cu102
  • Datasets 1.9.0
  • Tokenizers 0.10.3

Configuration

Architecture
Wav2Vec2ForSequenceClassification
Layers
24
Hidden size
1,024
Feed-forward size
4,096
Attention heads
16
Vocabulary size
33
Model type
wav2vec2

Identity and Version

Repository
ehcalabres/wav2vec2-lg-xlsr-en-speech-emotion-recognition
Publisher
Enrique Hernández Calabrés
Task
Audio classification
Modality
Audio
Library
transformers
Parameters
316M parameters
Languages
Not stated by the source
Revision
b520c9c46a719e36e1b9a91cad2cb5d0668757d8
First published
2022-03-02
Last updated
2024-10-24

Files and Weights

14 files, 2.5 GB in total. The weights are 3 files totalling 2.5 GB in bin, safetensors.

Weights3 files · 2.5 GB
Configuration2 files · 2.5 KB
Documentation1 file · 3.0 KB
Other6 files · 42.4 KB
Repository2 files · 804 B
Every file
FileTypeSizeSHA-256
model.safetensorsWeights1.3 GB 33bd858e5a4d
pytorch_model.binWeights1.3 GB 51f2f61d0b1e
training_args.binWeights2.9 KB 73d0e655c82c
config.jsonConfiguration2.3 KB
preprocessor_config.jsonConfiguration214 B
README.mdDocumentation3.0 KB
runs/Jul14_08-52-02_ea6be2bf8cd5/1626253006.9531522/events.out.tfevents.1626253006.ea6be2bf8cd5.900.5Other4.6 KB 460b2ae294cc
runs/Jul14_08-52-02_ea6be2bf8cd5/events.out.tfevents.1626253006.ea6be2bf8cd5.900.4Other5.0 KB af1accc9a006
runs/Jul14_08-58-15_ea6be2bf8cd5/1626253103.0537474/events.out.tfevents.1626253103.ea6be2bf8cd5.1946.1Other4.6 KB 369f6ff7d317
runs/Jul14_08-58-15_ea6be2bf8cd5/events.out.tfevents.1626253103.ea6be2bf8cd5.1946.0Other11.9 KB 9350f7355b22
runs/Sep21_17-17-21_cc179d296121/1632245116.2700055/events.out.tfevents.1632245116.cc179d296121.76.1Other4.5 KB bbb6d183f8cb
runs/Sep21_17-17-21_cc179d296121/events.out.tfevents.1632245116.cc179d296121.76.0Other11.9 KB fdd4dc2a0c25
.gitattributesRepository791 B
.gitignoreRepository13 B

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
2.5 GB
Download from Enrique Hernández Calabrés

Released by Enrique Hernández Calabrés through its official repository on Hugging Face. Read the license.

Memory Requirements

PrecisionWeights in memory
As published2.5 GB
16-bit0.6 GB
8-bit0.3 GB
4-bit0.2 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About wav2vec2-lg-xlsr-en-speech-emotion-recognition

How much GPU memory does wav2vec2-lg-xlsr-en-speech-emotion-recognition need?

About 0.8 GB at 16-bit and 0.2 GB at 4-bit: the weights (316M parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run wav2vec2-lg-xlsr-en-speech-emotion-recognition on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use wav2vec2-lg-xlsr-en-speech-emotion-recognition commercially?

Yes. wav2vec2-lg-xlsr-en-speech-emotion-recognition is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Similar Models

Speech emotion recognition for Russian over seven classes: anger, disgust, enthusiasm, fear, happiness, neutral, sadness. Fine-tuned from jonatasgrosman/wav2vec2-large-xlsr-53-russian on Aniemore/resd. Audio resampled to 16 kHz mono, clips capped at 12 s, normalized per utterance, padding masked. UA is macro-averaged recall, WA is accuracy, F1 is macro-averaged. All three test sets went through the same harness, so the rows are comparable to each other. The RESD split matches fold 1 of EmoBox bit for bit. The top entry there is WavLM-large at WA 56.47 / UA 55.87 / F1 55.82. These numbers are higher, but the training protocol differs — EmoBox freezes the encoder and trains a probe, this is a…

Open weights mit 316M parameters transformers

Model · Audio classification

wavlm-emotion-russian-resd

Aniemore

Speech emotion recognition for Russian over seven classes: anger, disgust, enthusiasm, fear, happiness, neutral, sadness. Fine-tuned from jonatasgrosman/expw2v2truwavlms363 on Aniemore/resd. Audio resampled to 16 kHz mono, clips capped at 12 s, normalized per utterance, padding masked. UA is macro-averaged recall, WA is accuracy, F1 is macro-averaged. All three test sets went through the same harness, so the rows are comparable to each other. The RESD split matches fold 1 of EmoBox bit for bit. The top entry there is WavLM-large at WA 56.47 / UA 55.87 / F1 55.82. These numbers are higher, but the training protocol differs — EmoBox freezes the encoder and trains a probe, this is a full…

Open weights mit 317M parameters transformers

The pre-trained model is this one - facebook/hubert-large-ls960-ft The DUSHA dataset used can be found here Fine-tuned in Google Colab using Pro account with A100 GPU Freezed all layers exept projector, classifier and all 24 HubertEncoderLayerStableLayerNorm layers Used half of the train dataset - 2 epochs - train batch size = 8 - eval batch size = 8 - gradient accumulation steps = 4 - learning rate = 5e-5 without warm up and decay Achieved - accuracy = 0.86 - balanced = 0.76 - macro f1 score = 0.81 on test set, improving accucary and f1 score compared to dataset baseline

Open weights apache-2.0 316M parameters transformers

This model is a fine-tuned version of facebook/wav2vec2-xls-r-300m on Librispeech-clean-100 for gender recognition. It achieves the following results on the evaluation set: The Librispeech-clean-100 dataset was used to train the model, with 70% of the data used for training, 10% for validation, and 20% for testing. The following hyperparameters were used during training: - learningrate: 3e-05 - trainbatchsize: 4 - evalbatchsize: 4 - gradientaccumulationsteps: 4 - totaltrainbatchsize: 16 - lrschedulertype: linear - lrschedulerwarmupratio: 0.1 - numepochs: 1 - mixedprecisiontraining: Native AMP - Transformers 4.28.0 - Pytorch 2.0.0+cu118 - Tokenizers 0.13.3

Open weights apache-2.0 316M parameters transformers

Model · Audio classification

wav2vec-vm-finetune

Jake Downie

This model is a fine-tuned version of facebook/wav2vec2-xls-r-300m for voicemail detection. It is trained on a dataset of call recordings to distinguish between voicemail greetings and live human responses. This model builds on wav2vec2-xls-r-300m, a self-supervised speech model trained on large-scale multilingual data. We fine-tuned it on the first two seconds of a call. - Automated voicemail detection in AI-powered call assistants. - Filtering voicemail responses in customer service and sales call automation. - Only trianed on the English language. - Assumes the voicemail track is isolated and contains no audio from the caller. - Designed for the first two seconds of audio when calling a…

Open weights apache-2.0 316M parameters transformers