Human-side speech from production call recordings, cut into utterance-level chunks by a two-engine VAD (Silero + TEN) and transcribed by third-party ASR providers. Each row keeps the transcript, the provider's confidence, and full provenance back to the source recording. One config per transcription system, so their output stays separable. The combined config interleaves several transcription systems within each shard, so this is the breakdown across the whole dataset. The systems differ substantially, so treat them as separate sources when training. Each row carries the system that produced it in its provider and model columns; the labels below are withheld aliases for the same systems, in…
Access requested at publisher
Chaashini — Hindi/Urdu for sugar syrup — is a continuously growing corpus of clean, single-speaker, studio-grade Indian-language speech built for training speech models (text-to-speech, speech recognition, speech language models). Every clip in the corpus has passed a strict multi-stage quality gate; the aim is purity over volume. The corpus grows automatically: new shards are appended every ~2 hours of newly accepted audio. Audio is sourced from publicly available spoken-word recordings (talks, interviews, narration, lectures, podcasts and similar long-form speech). Each recording then passes through: 1. Source-level screening – recordings dominated by music, singing, or non-speech content…
Access requested at publisher
apache-2.0
1M<n<10M