SAVRN
Search Contact SAVRN

Organization

VISTEC-depa AI Research Institute of Thailand

airesearch

Models in Library1
Datasets in Library0
Models on Hugging Face200
Followers116

Models

Finetuning wav2vec2-large-xlsr-53 on Thai Common Voice 7.0 We finetune wav2vec2-large-xlsr-53 based on Fine-tuning Wav2Vec2 for English ASR using Thai examples of Common Voice Corpus 7.0. The notebooks and scripts can be found in vistec-ai/wav2vec2-large-xlsr-53-th. The pretrained model and processor can be found at airesearch/wav2vec2-large-xlsr-53-th. Add syllabletokenize, wordtokenize (PyThaiNLP) and deepcut tokenizers to eval.py from robust-speech-event Common Voice Corpus 7.0](https://commonvoice.mozilla.org/en/datasets) contains 133 validated hours of Thai (255 total hours) at 5GB. We pre-tokenize with pythainlp.tokenize.wordtokenize. We preprocess the dataset using cleaning rules…

Open weights cc-by-sa-4.0 transformers