SAVRN
Search Contact SAVRN

Open-weight model · Speech recognition

wav2vec2-xls-r-300m-bengali

by Arijit arijitx/wav2vec2-xls-r-300m-bengali

This model is a fine-tuned version of facebook/wav2vec2-xls-r-300m on the OPENSLRSLR53 - bengali dataset. It achieves the following results on the evaluation set.

Parameters
Context
Weights4.7 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads1.1M

Model Card

By Arijit, published under apache-2.0, revision 45ed7c704f27.

This model is a fine-tuned version of facebook/wav2vec2-xls-r-300m on the OPENSLRSLR53 - bengali dataset. It achieves the following results on the evaluation set. With 5 gram language model trained on 30M sentences randomly chosen from AI4Bharat IndicCorp dataset: Note: 5% of a total 10935 samples have been used for evaluation. Evaluation set has 10935 examples which was not part of training training was done on first 95% and eval was done on last 5%. Training was stopped after 180k steps. Output predictions are available under files section. The following hyperparameters were used during training: - datasetname="openslr" - modelnameorpath="facebook/wav2vec2-xls-r-300m"…

Read Arijit's full model card

This model is a fine-tuned version of facebook/wav2vec2-xls-r-300m on the OPENSLR_SLR53 - bengali dataset. It achieves the following results on the evaluation set.

Without language model : - WER: 0.21726385291857586 - CER: 0.04725010353701041

With 5 gram language model trained on 30M sentences randomly chosen from AI4Bharat IndicCorp dataset : - WER: 0.15322879016421437 - CER: 0.03413696666806267

Note : 5% of a total 10935 samples have been used for evaluation. Evaluation set has 10935 examples which was not part of training training was done on first 95% and eval was done on last 5%. Training was stopped after 180k steps. Output predictions are available under files section.

Training hyperparameters

The following hyperparameters were used during training:

  • dataset_name="openslr"
  • model_name_or_path="facebook/wav2vec2-xls-r-300m"
  • dataset_config_name="SLR53"
  • output_dir="./wav2vec2-xls-r-300m-bengali"
  • overwrite_output_dir
  • num_train_epochs="50"
  • per_device_train_batch_size="32"
  • per_device_eval_batch_size="32"
  • gradient_accumulation_steps="1"
  • learning_rate="7.5e-5"
  • warmup_steps="2000"
  • length_column_name="input_length"
  • evaluation_strategy="steps"
  • text_column_name="sentence"
  • chars_to_ignore , ? . ! - \; \: \" “ % ‘ ” � — ’ … –
  • save_steps="2000"
  • eval_steps="3000"
  • logging_steps="100"
  • layerdrop="0.0"
  • activation_dropout="0.1"
  • save_total_limit="3"
  • freeze_feature_encoder
  • feat_proj_dropout="0.0"
  • mask_time_prob="0.75"
  • mask_time_length="10"
  • mask_feature_prob="0.25"
  • mask_feature_length="64"
  • preprocessing_num_workers 32

Framework versions

  • Transformers 4.16.0.dev0
  • Pytorch 1.10.1+cu102
  • Datasets 1.17.1.dev0
  • Tokenizers 0.11.0

Notes - Training and eval code modified from : https://github.com/huggingface/transformers/tree/master/examples/research_projects/robust-speech-event. - Bengali speech data was not available from common voice or librispeech multilingual datasets, so OpenSLR53 has been used. - Minimum audio duration of 0.5s has been used to filter the training data which excluded may be 10-20 samples. - OpenSLR53 transcripts are not part of LM training and LM used to evaluate.

Configuration

Architecture
Wav2Vec2ForCTC
Layers
24
Hidden size
1,024
Feed-forward size
4,096
Attention heads
16
Vocabulary size
112
Stored precision
float32
Model type
wav2vec2

Identity and Version

Repository
arijitx/wav2vec2-xls-r-300m-bengali
Publisher
Arijit
Task
Speech recognition
Modality
Audio
Library
transformers
Parameters
Not stated by the source
Languages
bn
Revision
45ed7c704f276acb9ed6ef234b66e67a1a2cb864
First published
2022-03-02
Last updated
2022-03-23

Files and Weights

37 files, 4.8 GB in total. The weights are 3 files totalling 4.7 GB in bin.

Weights3 files · 4.7 GB
Configuration15 files · 48.6 KB
Tokenizer2 files · 1.4 KB
Documentation2 files · 3.6 KB
Other13 files · 61.2 MB
Repository2 files · 1.3 KB
Every file
FileTypeSizeSHA-256
language_model/5gram.binWeights3.5 GB 59aa3014b4ee
pytorch_model.binWeights1.3 GB 0074e749c9ca
training_args.binWeights3.1 KB 6d386bae03a1
.ipynb_checkpoints/added_tokens-checkpoint.jsonConfiguration25 B
.ipynb_checkpoints/alphabet-checkpoint.jsonConfiguration755 B
.ipynb_checkpoints/eval-checkpoint.pyConfiguration5.7 KB
.ipynb_checkpoints/preprocessor_config-checkpoint.jsonConfiguration261 B
.ipynb_checkpoints/special_tokens_map-checkpoint.jsonConfiguration309 B
.ipynb_checkpoints/vocab-checkpoint.jsonConfiguration1.2 KB
added_tokens.jsonConfiguration25 B
alphabet.jsonConfiguration755 B
config.jsonConfiguration2.0 KB
eval.pyConfiguration5.7 KB
language_model/.ipynb_checkpoints/attrs-checkpoint.jsonConfiguration78 B
language_model/attrs.jsonConfiguration78 B
preprocessor_config.jsonConfiguration261 B
run_speech_recognition_ctc.pyConfiguration31.1 KB
special_tokens_map.jsonConfiguration309 B
.ipynb_checkpoints/README-checkpoint.mdDocumentation476 B
README.mdDocumentation3.1 KB
.ipynb_checkpoints/eval_run-checkpoint.shOther92 B
eval_run.shOther92 B
evaluation_no_lm/log_openslr_SLR53_train[95]_predictions.txtOther655.8 KB
evaluation_no_lm/log_openslr_SLR53_train[95]_targets.txtOther661.6 KB
evaluation_no_lm/openslr_SLR53_train[95]_eval_results.txtOther49 B
evaluation_with_lm/.ipynb_checkpoints/log_openslr_SLR53_train[95]_predictions-checkpoint.txtOther656.9 KB
evaluation_with_lm/.ipynb_checkpoints/log_openslr_SLR53_train[95]_targets-checkpoint.txtOther661.6 KB
evaluation_with_lm/.ipynb_checkpoints/openslr_SLR53_train[95]_eval_results-checkpoint.txtOther49 B
evaluation_with_lm/log_openslr_SLR53_train[95]_predictions.txtOther656.9 KB
evaluation_with_lm/log_openslr_SLR53_train[95]_targets.txtOther661.6 KB
evaluation_with_lm/openslr_SLR53_train[95]_eval_results.txtOther49 B
language_model/unigrams.txtOther57.2 MB f98e7b082bc4
run.shOther1.1 KB
.gitattributesRepository1.2 KB
.gitignoreRepository13 B
tokenizer_config.jsonTokenizer287 B
vocab.jsonTokenizer1.2 KB

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
4.7 GB
Download from Arijit

Released by Arijit through its official repository on Hugging Face. Read the license.

Built From

  • Trained on (disclosed) AI4Bharat/IndicCorp
  • Trained on (disclosed) SLR53
  • Trained on (disclosed) openslr

Evaluations

Each result is shown as reported, with the conditions its reporter stated. None is a SAVRN measurement. A comparison lines two results up only when their configuration, unit and setup are all stated and identical.

BenchmarkConditionsResultReported byRevisionDate
Open SLR Task Speech RecognitionMetric Test CERComparison conditions not established 0.0472501 arijitx
Publisher reported
Evaluated revision not stated
Open SLR Task Speech RecognitionMetric Test CER with lmComparison conditions not established 0.034137 arijitx
Publisher reported
Evaluated revision not stated
Open SLR Task Speech RecognitionMetric Test WERComparison conditions not established 0.217264 arijitx
Publisher reported
Evaluated revision not stated
Open SLR Task Speech RecognitionMetric Test WER with lmComparison conditions not established 0.153229 arijitx
Publisher reported
Evaluated revision not stated

Memory Requirements

PrecisionWeights in memory
As published4.7 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About wav2vec2-xls-r-300m-bengali

Can I use wav2vec2-xls-r-300m-bengali commercially?

Yes. wav2vec2-xls-r-300m-bengali is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Similar Models

Model · Speech recognition

wav2vec2-large-xlsr-53-japanese

Jonatas Grosman

Fine-tuned facebook/wav2vec2-large-xlsr-53 on Japanese using the train and validation splits of Common Voice 6.1, CSS10 and JSUT. When using this model, make sure that your speech input is sampled at 16kHz. This model has been fine-tuned thanks to the GPU credits generously given by the OVHcloud:) The script used for training can be found here: https://github.com/jonatasgrosman/wav2vec2-sprint The model can be used directly (without a language model) as follows... Using the HuggingSound library: The model can be evaluated as follows on the Japanese test data of Common Voice. In the table below I report the Word Error Rate (WER) and the Character Error Rate (CER) of the model. I ran the…

Open weights apache-2.0 transformers

Model · Speech recognition

whisperkit-coreml

Argmax

WhisperKit is part of Argmax OSS, an On-device Speech AI SDK for Apple Silicon: https://github.com/argmaxinc/argmax-oss-swift Check out the WhisperKit paper and presentation from ICML 2025: https://icml.cc/virtual/2025/47854 For real-time transcription with speakers and custom vocabulary, check out Argmax Pro SDK: https://www.argmaxinc.com/blog/argmax-sdk-2

Open weights mit whisperkit

Model · Speech recognition

speaker-diarization-3.1

Pyannote

Using this open-source model in production? Consider switching to pyannoteAI for better and faster options. This pipeline is the same as pyannote/speaker-diarization-3.0 except it removes the problematic use of onnxruntime. Both speaker segmentation and embedding now run in pure PyTorch. This should ease deployment and possibly speed up inference. It requires pyannote.audio version 3.1 or higher. It ingests mono audio sampled at 16kHz and outputs speaker diarization as an Annotation instance: - stereo or multi-channel audio files are automatically downmixed to mono by averaging the channels. - audio files sampled at a different rate are resampled to 16kHz automatically upon loading. 1.…

Access requested at publisher mit pyannote-audio

Fine-tuned facebook/wav2vec2-large-xlsr-53 on Portuguese using the train and validation splits of Common Voice 6.1. When using this model, make sure that your speech input is sampled at 16kHz. This model has been fine-tuned thanks to the GPU credits generously given by the OVHcloud:) The script used for training can be found here: https://github.com/jonatasgrosman/wav2vec2-sprint The model can be used directly (without a language model) as follows... Using the HuggingSound library: 1. To evaluate on mozilla-foundation/commonvoice60 with split test 2. To evaluate on speech-recognition-community-v2/devdata If you want to cite this model you can use this

Open weights apache-2.0 transformers

Model · Speech recognition

speaker-diarization-community-1

Pyannote

This pipeline ingests mono audio sampled at 16kHz and outputs speaker diarization. - stereo or multi-channel audio files are automatically downmixed to mono by averaging the channels. - audio files sampled at a different rate are resampled to 16kHz automatically upon loading. The main improvements brought by Community-1 are: - improved speaker assignment and counting - simpler reconciliation with transcription timestamps with exclusive speaker diarization - easy offline use (i.e. without internet connection) - (optionally) hosted on pyannoteAI cloud 1. pip install pyannote.audio 3. Create access token at hf.co/settings/tokens. Out of the box, Community-1 is much better than…

Access requested at publisher cc-by-4.0 pyannote-audio

Model · Speech recognition

wav2vec2-large-xlsr-53-russian

Jonatas Grosman

Fine-tuned facebook/wav2vec2-large-xlsr-53 on Russian using the train and validation splits of Common Voice 6.1 and CSS10. When using this model, make sure that your speech input is sampled at 16kHz. This model has been fine-tuned thanks to the GPU credits generously given by the OVHcloud:) The script used for training can be found here: https://github.com/jonatasgrosman/wav2vec2-sprint The model can be used directly (without a language model) as follows... Using the HuggingSound library: 1. To evaluate on mozilla-foundation/commonvoice60 with split test 2. To evaluate on speech-recognition-community-v2/devdata If you want to cite this model you can use this

Open weights apache-2.0 transformers