SAVRN
Search Contact SAVRN

whisper-small-malayalam · Model Card

whisper-small-malayalam: Model Card

Written by Sajil C K, published under apache-2.0, revision dbd72416303a, read 2026-10-09. Shown as written; SAVRN's own facts about this model are on its page.

Fine-tuned version of openai/whisper-small on a multi-corpus Malayalam speech dataset.

Model Description

  • Base model: openai/whisper-small (244M parameters)
  • Language: Malayalam (ml)
  • Task: Automatic Speech Recognition (transcription)
  • Training steps: 3500
  • Best WER: 37.64% on CommonVoice 25 Malayalam test set
  • Leakage-free multi-source test: 47.8% WER / 13.6% CER (231 clips across 6 sources, none with a transcript in the training split — see Multi-source evaluation)
  • CPU speed (Transformers, FP32, 4 vCPU): RTF 1.96, i.e. slower than real time. For CPU deployment use the whisper.cpp builds

Training Data

The model was trained on an aggregated corpus of 5 Malayalam speech datasets, combined and published as sajilck/malayalam-asr-corpus.

Corpus Source Domain Access
IMaSC thennal/imasc TTS / Read speech HuggingFace
SMC Malayalam Speech Corpus sajilck/smc-malayalam-speech-corpus Read speech Kaggle
IndicTTS Malayalam kavyamanohar/indic-tts-malayalam-speech-corpus TTS / Read speech Kaggle
OpenSLR 63 sajilck/openslr63 Crowdsourced Kaggle
CommonVoice 25 Malayalam sajilck/common-voice-malayalam Crowdsourced Kaggle

Total: ~86,000 samples across TTS-recorded, read speech, and crowdsourced domains.

Benchmark Results

Model Params WER ↓ Notes
openai/whisper-small (base) 244M ~85% No Malayalam fine-tuning
smcproject/Malwhisper-v1-medium 769M 61.84% Single corpus (IMaSC only)
sajilck/whisper-small-malayalam 244M 37.64% Multi-corpus fine-tuning

Key advantages over prior work: - 3× smaller model than Malwhisper-v1-medium, better WER - 5 corpora vs 1 — better speaker and domain diversity - Multi-domain training — TTS, read speech, and crowdsourced audio

Multi-source evaluation

The CommonVoice figure above covers one domain. To see how the model does across all domains, it was scored on the corpus's leakage-free held-out set: the 2,684 test rows whose transcript does not appear, exactly or as a near-duplicate, in the train split (see Limitations). Up to 50 clips per source were drawn at random (231 clips, 18.6 min of audio; some sources have fewer clean rows than that) and decoded with FP32 greedy decoding. Evaluated October 2026.

Metric Result
WER, normalised (95% bootstrap CI) 47.8% (44.0% – 51.6%)
WER, raw 49.2%
CER, normalised 13.6%
Substitutions / deletions / insertions 32.4% / 10.4% / 5.0%
Utterances fully correct 15.2% (20.3% if word boundaries are ignored)
Empty outputs / repetition loops 0 / 0
Source Clips WER CER
IndicTTS 50 37.0% 9.5%
OpenSLR 63 32 41.0% 6.1%
IMaSC 44 42.5% 8.0%
Shrutilipi 50 56.8% 22.3%
CommonVoice 50 58.1% 17.2%
SMC 5 35.0% 3.9%

SMC has only 5 clean test rows, so its score is not meaningful on its own.

Normalisation steps: Unicode NFC, old-style chillu (consonant + virama + ZWJ) converted to atomic chillu, zero-width joiners removed, and punctuation stripped.

How to read these numbers

  • The headline is source-balanced. It weights each source roughly equally. The clean held-out set itself is 88% Shrutilipi; weighted by that mix, WER would be about 55% (an estimate from the per-source scores).
  • Clean vs leaked rows. On 60 test rows whose transcripts do appear in training, WER was 32.5%, against 47.8% on clean rows. Part of that gap is memorisation, but part is source mix: the leaked rows are mostly IMaSC, which is easier read speech. The like-for-like comparison is IMaSC on its own, which went from 23.9% on the leaky split to 42.5% on clean rows.
  • WER is high partly because Malayalam words are long. One wrong character makes a whole agglutinated word wrong, so CER (13.6%) is the better guide to how much of the speech is recognised. Word-boundary differences (മുഖ്യമന്ത്രി vs മുഖ്യ മന്ത്രി) account for only a small share: ignoring spaces raises fully-correct utterances from 15.2% to 20.3%. Most errors are real character-level mistakes, such as കുടുംബം → ഉടുംബം.
  • The 58.1% CommonVoice figure does not replace 37.64%. It comes from 50 CommonVoice-sourced rows of this corpus's own split, with a wide confidence interval and different normalisation. It is not the official CommonVoice 25 test set.

An earlier run (September 2026, 300 clips) drew from the full test split, including leaked rows, and scored 42.9% WER / 12.3% CER. Those numbers are superseded by the leakage-free results above.

Limitations

This section is maintained openly and updated as new evaluation work uncovers issues. Last updated after the leakage-free re-evaluation (October 2026).

Original benchmark was narrow. The initially reported 37.64% WER was measured only on the CommonVoice Malayalam test split — read, studio-quality speech. It was not evaluated against the other four source domains in the training corpus (IMaSC, SMC, IndicTTS, OpenSLR 63) or against broadcast/radio-news-style audio.

The corpus's official train/test split has confirmed leakage. An exact-transcript-hash and fuzzy near-duplicate check between this model's training corpus's train (86,911 rows) and test (4,828 rows) splits found: - 41.6% of test rows have an exact-duplicate transcript in train - 44.4% are flagged by fuzzy near-duplicate matching

Any WER computed on the raw test split partly reflects memorization, not generalization. On the leakage-free rows the model scores 47.8% WER, against 42.9% on the full split; IMaSC alone moves from 23.9% to 42.5%.

Leakage is concentrated in the smaller, fixed-script sources. Per-source kept rate after filtering out flagged rows:

Source Kept %
IMaSC 2.7%
SMC 5.7%
OpenSLR 63 13.9%
CommonVoice 70.7%
IndicTTS 84.6%
Shrutilipi 93.0%

IMaSC and SMC — small corpora with a bounded set of scripted sentences read by a handful of speakers — are almost entirely duplicated across the split. Shrutilipi, despite being scraped at document scale, is the cleanest source by this measure.

The clean held-out set is source-imbalanced. After filtering, 2,684 clean rows remain, of which 87.6% are Shrutilipi. Only 44 IMaSC, 32 OpenSLR 63 and 5 SMC rows survive, so per-source scores for those sources have wide error bars. The headline evaluation samples sources evenly to avoid being a Shrutilipi-only number.

The GGML/GGUF "WER gap" is mostly explained by the evaluation data. whisper.cpp f16/q5_0/q8_0 conversions scored 51.7%–54.4% WER on a 20-clip sample of the multi-source test split, far above the 37.64% CommonVoice figure. A later CPU evaluation of the original FP32 Transformers model on the same split scored 42.9% on 300 clips and 51.1% on a 24-clip subset. The unconverted model shows a gap of the same size, so most of it comes from the broader, harder source mix and the noise of small samples, not from the conversion. Quantization itself still costs a little accuracy: about 2.7 WER points from f16 to q8_0, and about 5.6 points for PyTorch INT8 dynamic quantization on the 24-clip subset.

What's fixed: a leakage-free held-out set exists, and the multi-source evaluation has been re-run on it (47.8% WER / 13.6% CER). Still open: source imbalance in that set (too few clean read-speech rows), and checking the train split for internal duplicates.

General limitations

  • Trained on read speech and crowdsourced audio, so it may perform worse on spontaneous, conversational Malayalam
  • Foreign proper nouns (English names, place names) may be transcribed with Malayalam phonetic approximations
  • Performance may vary across Malayalam dialects
  • Not real-time on CPU with plain Transformers (see CPU inference benchmark)

Full writeup with methodology: I Checked My Own ASR Dataset for Leakage — Here's What I Found

Usage

HuggingFace Transformers (Python)

from transformers import pipeline

pipe = pipeline(
    "automatic-speech-recognition",
    model="sajilck/whisper-small-malayalam",
    generate_kwargs={"language": "malayalam", "task": "transcribe"},
)

result = pipe("your_audio.wav")
print(result["text"])

For longer audio files:

pipe = pipeline(
    "automatic-speech-recognition",
    model="sajilck/whisper-small-malayalam",
    generate_kwargs={"language": "malayalam", "task": "transcribe"},
    chunk_length_s=30,
    stride_length_s=5,
)
result = pipe("long_audio.wav")
print(result["text"])

CPU Inference Benchmark (Transformers, no GPU)

This measures the model on CPU through plain HuggingFace Transformers with default settings. It is a baseline, not an optimised setup.

Setting Value
Hardware Kaggle CPU: Intel Xeon @ 2.20 GHz, 4 vCPU (2 physical cores), 33.7 GB RAM
Software PyTorch 2.10 (CPU), Transformers 5.0
Decoding FP32, greedy, batch size 1, max_new_tokens=225
Evaluation data 300 clips (average 4.8 s long) from the corpus test split, September 2026
Metric Result
Real-time factor (RTF, all clips together) 1.96 (0.51× real time)
RTF median / p90 per clip 2.05 / 2.97
Latency per clip, p50 / p90 / max 9.1 s / 14.1 s / 15.9 s
Cold start (first clip, 3.5 s long) 4.5 s
Decoding speed ~13 tokens/s (~77 ms per token)
Model load time 13.6 s
Peak RAM ~3.9 GB

RTF = processing time ÷ audio duration. Values above 1 mean slower than real time. The October leakage-free run on the same hardware gave consistent figures (RTF 2.17, p90 latency 15.3 s, 12.2 tokens/s).

Where the time goes. Feature extraction takes about 0.1% of the time; almost all of it is the decoder generating text one token at a time. Whisper's tokenizer splits Malayalam script into many small pieces: this run produced about 120 tokens per clip, roughly 25 tokens per second of audio. Speed on CPU is therefore limited mainly by transcript length in tokens, not by audio length.

Clip length Clips RTF p90 latency
< 3 s 67 2.60 8.3 s
3–6 s 169 2.09 12.1 s
6–10 s 54 1.72 14.7 s
10–15 s 9 1.16 14.9 s

Settings for faster inference (24-clip subset; compare these rows with each other only)

Setting WER RTF p90 latency Size
FP32, greedy (default) 51.1% 1.80 14.5 s 967 MB
INT8 dynamic quantisation (torch.ao), greedy 56.7% 1.40 11.4 s 413 MB
FP32, beam search (5 beams) 51.7% 4.80 41.1 s 967 MB
CPU threads RTF Speed-up vs 1 thread
1 3.18 1.00×
2 1.83 1.73×
4 1.81 1.76×
  • Use greedy decoding. Beam search with 5 beams was about 2.7× slower and no more accurate.
  • Set threads to the number of physical cores. Hyperthreads added almost nothing (2 → 4 threads: +1%).
  • Real-time streaming isn't feasible with this setup. With 6 s windows and a 4 s hop, p90 latency per window was 12.9 s, about 3.2× over budget. For CPU use, run the whisper.cpp builds below. The RTFs in that table were measured on 30 s audio, so they aren't directly comparable with the numbers here.

Evaluation notebooks: whisper-malayalam-cpu-eval.ipynb (speed, September 2026) and whisper-malayalam-clean-eval.ipynb (leakage-free accuracy, October 2026), both Kaggle CPU only.

GGML / whisper.cpp — CPU Inference (No GPU Required)

Quantized GGML variants are available for use with whisper.cpp, enabling Malayalam ASR on any laptop or edge device without a GPU.

Variant File Size RTF (30s audio) Speed Notes
FP16 ggml-model-f16.bin 487 MB 0.40 2.6× real-time Best quality
Q5_0 ggml-model-q5_0.bin 175 MB 0.44 2.3× real-time Smallest size
Q8_0 ggml-model-q8_0.bin 264 MB 0.34 3.0× real-time Recommended

Benchmarked on Kaggle CPU (4 cores, AVX2). Q8_0 outperforms F16 and Q5_0 due to efficient SIMD integer operations on AVX2 hardware. WER difference between F16 and Q8_0 is only 2.68%, making Q8_0 the best overall choice.

Note: RTF is high for short clips (<10s) due to fixed 30s mel spectrogram encoding overhead. All variants process 30s+ audio faster than real-time.

Quick Start

# 1. Build whisper.cpp
git clone https://github.com/ggml-org/whisper.cpp
cd whisper.cpp
cmake -B build && cmake --build build --config Release

# 2. Download Q8_0 (recommended)
wget https://huggingface.co/sajilck/whisper-small-malayalam/resolve/main/ggml/ggml-model-q8_0.bin

# 3. Transcribe Malayalam audio (16kHz mono WAV)
./build/bin/whisper-cli -m ggml-model-q8_0.bin -l ml -f your_audio.wav

Audio Requirements

  • Format: WAV (16-bit PCM)
  • Sample rate: 16 kHz
  • Channels: Mono

Convert any audio to the required format:

ffmpeg -i input.mp3 -ar 16000 -ac 1 -c:a pcm_s16le output.wav

Training Details

Parameter Value
Base model openai/whisper-small
Training steps 3500
Effective batch size 16 (batch=4, grad_accum=4)
Learning rate 1e-5
Warmup steps 500
Precision fp16
Hardware NVIDIA Tesla P100 16GB
Framework HuggingFace Transformers + Seq2SeqTrainer

Training Data Preprocessing

Audio from all 5 corpora was: - Resampled to 16kHz mono - Filtered to 0.5–30 second clips - Converted to log-mel spectrograms (80 mel bins) - Tokenized using Whisper's multilingual tokenizer with language token <|ml|>

Future Work

  • v2: Adding Shrutilipi broadcast news corpus with higher learning rate (lr=2e-4) based on findings from Adalat AI Vividh-ASR paper
  • whisper-medium-malayalam and whisper-tiny-malayalam variants
  • Evaluation on Vividh-ASR benchmark across all 4 speech difficulty tiers
  • Expand the clean held-out set for read-speech sources (IMaSC, SMC, OpenSLR 63)
  • Faster CPU inference for real-time streaming: benchmark CTranslate2 / faster-whisper INT8 and whisper.cpp on the same clips, targeting a 6 s window in under 4 s

Citation

@misc{sajilck2026whispermalayalam,
  author = {Sajil C.K.},
  title = {Whisper Small Malayalam: Multi-Corpus Fine-Tuning of Whisper for Malayalam ASR},
  year = {2026},
  publisher = {HuggingFace},
  url = {https://huggingface.co/sajilck/whisper-small-malayalam}
}

License

This model is released under the Apache 2.0 license, consistent with the base Whisper model. Training corpora retain their individual licenses — please refer to each source dataset for usage terms.