Fine-tuned facebook/wav2vec2-large-xlsr-53 on Telugu using the OpenSLR SLR66 dataset. When using this model, make sure that your speech input is sampled at 16kHz. The model can be used directly (without a language model) as follows: 70% of the OpenSLR Telugu dataset was used for training. Train Split of annotations is here Test Split of annotations is here Training Data Preparation notebook can be found here Training notebook can be foundhere Evaluation notebook is here