SAVRN
Search Contact SAVRN

SAVRN Model Hub · Datasets by Task

Audio Classification Datasets

5 open-weight audio classification datasets in the SAVRN Model Hub, with Silencio Voice AI, Mandip Goswami and Alosh Denny publishing the most.

5Datasets
3Publishers
3Licenses

Most Downloaded

DatasetPublisherLicenseMonthly downloads
t2a-daddy Alosh Denny apache-2.0 3.3k
english-accents-speech Silencio Voice AI cc-by-nc-4.0 —
spanish-accents-speech Silencio Voice AI cc-by-nc-4.0 —
french-accents-speech Silencio Voice AI cc-by-nc-4.0 —
audio-eval-suite Mandip Goswami cc-by-4.0 —

Licenses

LicenseDatasetsCommercial use
cc-by-nc-4.03Not without separate permission
cc-by-4.01Yes
apache-2.01Yes

Who Publishes Them

PublisherDatasets
Silencio Voice AI3
Mandip Goswami1
Alosh Denny1

All 5 Datasets

Dataset · Audio classification

t2a-daddy

Alosh Denny

t2a-daddy Male-voice ASMR corpus for the text2asmr project. Previously published as aoxo/audios3. aoxo/t2a-audios-v1 (the original v1 corpus). Layout path what source audio, 48 kHz AAC, one folder per creator word-level Whisper large-v3 alignment ([] = skipped: near-silent or undecodable) labels/qwen3omni.jsonl non-speech ontology labels for gap clips

Publicly accessible apache-2.0

Dataset · Audio classification

audio-eval-suite

Mandip Goswami

A small, standardized audio-robustness evaluation set packaged to drop into eval loops. It pairs fully synthetic, license-clean speech-like probe signals with 300 real-synthetic room impulse responses (RIRs) and a labeled degradation chain (reverberation + additive noise), so you can measure how a model's acoustic predictions hold up as rooms get more reverberant and noisier. Think of it as a quick sanity/robustness check: one line to load, a concrete metric, and ground-truth room acoustics for every clip. Measured ground-truth ranges (from the underlying RIRs): - reverbbin distribution: mild 875, strong 600, clean 25 Sources are synthetic speech-LIKE probe signals, not real speech. Each is…

Publicly accessible cc-by-4.0 1K<n<10K

Dataset · Audio classification

english-accents-speech

Silencio Voice AI

English from 282 speakers born in 35 countries. 513 clips, 9.3 hours, each labelled with the speaker's country of birth, first language, self-reported accent or regional variety, declared English proficiency, age band, gender and recording device. Audio and speaker metadata only. Human-validated transcription is available on request. English as it is actually spoken worldwide is mostly not American or British. Across Silencio's English catalogue, 59% of recorded hours come from speakers born in Africa and only about 11% from the US, UK, Canada, Australia, Ireland and New Zealand combined. Nigeria, the largest single origin, is 16% of hours. This sample covers 35 countries of birth.…

Publicly accessible cc-by-nc-4.0 n<1K

Dataset · Audio classification

french-accents-speech

Silencio Voice AI

French from 20 speakers born in 9 countries. 32 clips, 0.5 hours, each labelled with the speaker's country of birth, first language, self-reported accent or regional variety, declared French proficiency, age band, gender and recording device. Audio and speaker metadata only. Human-validated transcription is available on request. Most French speakers today live in Africa, and Silencio's French catalogue reflects that: 74% of recorded hours come from speakers born in Africa, led by Benin, Senegal, Cameroon, Morocco and Madagascar. France, Belgium, Switzerland and Canada together are 21%. This sample covers 9 countries of birth, led by Senegal and Tunisia. 17 of the 20 speakers use French as a…

Publicly accessible cc-by-nc-4.0 n<1K

Dataset · Audio classification

spanish-accents-speech

Silencio Voice AI

Spanish from 12 speakers born in 9 countries. 19 clips, 0.3 hours, each labelled with the speaker's country of birth, first language, self-reported accent or regional variety, declared Spanish proficiency, age band, gender and recording device. Audio and speaker metadata only. Human-validated transcription is available on request. Silencio's Spanish catalogue is 65% Latin American by recorded hours, led by Venezuela, with Spain at 9%. It also holds a substantial body of Spanish spoken as a second language: 18% of hours come from speakers born in Africa, notably Nigeria, Kenya and Egypt. This sample covers 9 countries of birth across Latin America, Spain, Africa, the Middle East and the…

Publicly accessible cc-by-nc-4.0 n<1K

Other Tasks

See all