SAVRN
Search Contact SAVRN

SAVRN Model Hub · Datasets by Task

Text to speech Datasets

3 datasets in the SAVRN Model Hub for text to speech, from publishers including Yoach Lacombe, VoiceHub, Kapture CX.

3 datasets.

Dataset · Text to speech

cml-tts

Yoach Lacombe

CML-TTS is a recursive acronym for CML-Multi-Lingual-TTS, a Text-to-Speech (TTS) dataset developed at the Center of Excellence in Artificial Intelligence (CEIA) of the Federal University of Goias (UFG). CML-TTS is a dataset comprising audiobooks sourced from the public domain books of Project Gutenberg, read by volunteers from the LibriVox project. The dataset includes recordings in Dutch, German, French, Italian, Polish, Portuguese, and Spanish, all at a sampling rate of 24kHz. The data archives were restructured from the original ones from OpenSLR to make it easier to stream. - text-to-speech, text-to-audio: The dataset can also be used to train a model for Text-To-Speech (TTS). The…

Publicly accessible cc-by-4.0 1M<n<10M

Dataset · Text to speech

voicehub-arena-seed-tts-eval

VoiceHub

Incrementally published generated audio and WER, CER, DNSMOS, WavLM-large ECAPA speaker SIM and UTMOS22 measurements. The full campaign is still running. Each generation method is evaluated separately using its publisher's native API. Full evaluations contain all 1,088 English Seed-TTS-Eval targets; eight-target diagnostic pilots are stored separately and must not be treated as full scores. experiments/ / / / contains result.json, records.json, contract.json, verification.json, and audio.tar. records.json contains a rows array. Each successful row specifies its WAV's audioarchive, audiooffset, audiobytes, and audiosha256. Request exactly that byte range from the archive at an immutable…

Publicly accessible

Dataset · Text to speech

Chaashini

Kapture CX

Chaashini — Hindi/Urdu for sugar syrup — is a continuously growing corpus of clean, single-speaker, studio-grade Indian-language speech built for training speech models (text-to-speech, speech recognition, speech language models). Every clip in the corpus has passed a strict multi-stage quality gate; the aim is purity over volume. The corpus grows automatically: new shards are appended every ~2 hours of newly accepted audio. Audio is sourced from publicly available spoken-word recordings (talks, interviews, narration, lectures, podcasts and similar long-form speech). Each recording then passes through: 1. Source-level screening – recordings dominated by music, singing, or non-speech content…

Access requested at publisher apache-2.0 1M<n<10M

Who Publishes These Datasets

Other tasks

See all