Fine-tuned version of openai/whisper-small on a multi-corpus Malayalam speech dataset. - CPU speed (Transformers, FP32, 4 vCPU): RTF 1.96, i.e. slower than real time. For CPU deployment use the whisper.cpp builds The model was trained on an aggregated corpus of 5 Malayalam speech datasets, combined and published as sajilck/malayalam-asr-corpus. Total: ~86,000 samples across TTS-recorded, read speech, and crowdsourced domains. - 3× smaller model than Malwhisper-v1-medium, better WER - 5 corpora vs 1 — better speaker and domain diversity - Multi-domain training — TTS, read speech, and crowdsourced audio The CommonVoice figure above covers one domain. To see how the model does across all…