# Audio Classification Datasets: AI Datasets
Source: https://savrn.com/datasets/tasks/audio-classification
Markdown alternate of the page above; the site index is https://savrn.com/llms.txt

---

SAVRN Model Hub · Datasets by Task

# Audio Classification Datasets

6 open-weight audio classification datasets in the SAVRN Model Hub, with Silencio Voice AI, Mandip Goswami and CSY20060325 publishing the most.

6Datasets

4Publishers

4Licenses

## Most Downloaded

| Dataset | Publisher | License | Monthly downloads |
| --- | --- | --- | --- |
| [t2a-daddy](https://savrn.com/datasets/t2a-daddy) | Alosh Denny | apache-2.0 | 3.3k |
| [english-accents-speech](https://savrn.com/datasets/english-accents-speech) | Silencio Voice AI | cc-by-nc-4.0 | — |
| [spanish-accents-speech](https://savrn.com/datasets/spanish-accents-speech) | Silencio Voice AI | cc-by-nc-4.0 | — |
| [french-accents-speech](https://savrn.com/datasets/french-accents-speech) | Silencio Voice AI | cc-by-nc-4.0 | — |
| [audio-eval-suite](https://savrn.com/datasets/audio-eval-suite) | Mandip Goswami | cc-by-4.0 | — |
| [sjj](https://savrn.com/datasets/sjj) | CSY20060325 | Not stated | — |

## Licenses

| License | Datasets | Commercial use |
| --- | --- | --- |
| cc-by-nc-4.0 | 3 | Not without separate permission |
| not stated | 1 | Not stated |
| cc-by-4.0 | 1 | Yes |
| apache-2.0 | 1 | Yes |

## Who Publishes Them

| Publisher | Datasets |
| --- | --- |
| [Silencio Voice AI](https://savrn.com/model-publishers/silencionetwork) | 3 |
| [Mandip Goswami](https://savrn.com/model-publishers/mandipgoswami) | 1 |
| [CSY20060325](https://savrn.com/model-publishers/csy20060325) | 1 |
| [Alosh Denny](https://savrn.com/model-publishers/aoxo) | 1 |

## All 6 Datasets

Dataset · Audio classification

### [t2a-daddy](https://savrn.com/datasets/t2a-daddy)

[Alosh Denny](https://savrn.com/model-publishers/aoxo)

t2a-daddy Male-voice ASMR corpus for the text2asmr project. Previously published as aoxo/audios3. aoxo/t2a-audios-v1 (the original v1 corpus). Layout path what source audio, 48 kHz AAC, one folder per creator word-level Whisper large-v3 alignment ([] = skipped: near-silent or undecodable) labels/qwen3omni.jsonl non-speech ontology labels for gap clips

Publicly accessible apache-2.0

[View dataset](https://savrn.com/datasets/t2a-daddy)

Dataset · Audio classification

### [audio-eval-suite](https://savrn.com/datasets/audio-eval-suite)

[Mandip Goswami](https://savrn.com/model-publishers/mandipgoswami)

A small, standardized audio-robustness evaluation set packaged to drop into eval loops. It pairs fully synthetic, license-clean speech-like probe signals with 300 real-synthetic room impulse responses (RIRs) and a labeled degradation chain (reverberation + additive noise), so you can measure how a model's acoustic predictions hold up as rooms get more reverberant and noisier. Think of it as a quick sanity/robustness check: one line to load, a concrete metric, and ground-truth room acoustics for every clip. Measured ground-truth ranges (from the underlying RIRs): - reverbbin distribution: mild 875, strong 600, clean 25 Sources are synthetic speech-LIKE probe signals, not real speech. Each is…

Publicly accessible cc-by-4.0 1K<n<10K

[View dataset](https://savrn.com/datasets/audio-eval-suite)

Dataset · Audio classification

### [english-accents-speech](https://savrn.com/datasets/english-accents-speech)

[Silencio Voice AI](https://savrn.com/model-publishers/silencionetwork)

English from 282 speakers born in 35 countries. 513 clips, 9.3 hours, each labelled with the speaker's country of birth, first language, self-reported accent or regional variety, declared English proficiency, age band, gender and recording device. Audio and speaker metadata only. Human-validated transcription is available on request. English as it is actually spoken worldwide is mostly not American or British. Across Silencio's English catalogue, 59% of recorded hours come from speakers born in Africa and only about 11% from the US, UK, Canada, Australia, Ireland and New Zealand combined. Nigeria, the largest single origin, is 16% of hours. This sample covers 35 countries of birth.…

Publicly accessible cc-by-nc-4.0 n<1K

[View dataset](https://savrn.com/datasets/english-accents-speech)

Dataset · Audio classification

### [french-accents-speech](https://savrn.com/datasets/french-accents-speech)

[Silencio Voice AI](https://savrn.com/model-publishers/silencionetwork)

French from 20 speakers born in 9 countries. 32 clips, 0.5 hours, each labelled with the speaker's country of birth, first language, self-reported accent or regional variety, declared French proficiency, age band, gender and recording device. Audio and speaker metadata only. Human-validated transcription is available on request. Most French speakers today live in Africa, and Silencio's French catalogue reflects that: 74% of recorded hours come from speakers born in Africa, led by Benin, Senegal, Cameroon, Morocco and Madagascar. France, Belgium, Switzerland and Canada together are 21%. This sample covers 9 countries of birth, led by Senegal and Tunisia. 17 of the 20 speakers use French as a…

Publicly accessible cc-by-nc-4.0 n<1K

[View dataset](https://savrn.com/datasets/french-accents-speech)

Dataset · Audio classification

### [sjj](https://savrn.com/datasets/sjj)

[CSY20060325](https://savrn.com/model-publishers/csy20060325)

This release contains only selected synthetic training samples, the frozen L1/L2 full/hard evaluation sets, and the full merged V3 L1 calibrated250 checkpoint. The publisher has confirmed permission to redistribute the included data. Original source rights and attribution remain applicable; no blanket replacement license is asserted. - releases/20261004/train/: independently extractable shards, each with manifest.jsonl, selected audio, source hashes, and SHA256SUMS.json. Union membership is specified by selectedforl1 and selectedforl2; do not count a shared sample twice. No real training samples, rejected candidate regions, failed generations, or feature caches are included.…

Publicly accessible

[View dataset](https://savrn.com/datasets/sjj)

Dataset · Audio classification

### [spanish-accents-speech](https://savrn.com/datasets/spanish-accents-speech)

[Silencio Voice AI](https://savrn.com/model-publishers/silencionetwork)

Spanish from 12 speakers born in 9 countries. 19 clips, 0.3 hours, each labelled with the speaker's country of birth, first language, self-reported accent or regional variety, declared Spanish proficiency, age band, gender and recording device. Audio and speaker metadata only. Human-validated transcription is available on request. Silencio's Spanish catalogue is 65% Latin American by recorded hours, led by Venezuela, with Spain at 9%. It also holds a substantial body of Spanish spoken as a second language: 18% of hours come from speakers born in Africa, notably Nigeria, Kenya and Egypt. This sample covers 9 countries of birth across Latin America, Spain, Africa, the Middle East and the…

Publicly accessible cc-by-nc-4.0 n<1K

[View dataset](https://savrn.com/datasets/spanish-accents-speech)

## Other Tasks

- [Text Generation](https://savrn.com/datasets/tasks/text-generation) 96
- [Robotics](https://savrn.com/datasets/tasks/robotics) 82
- [Text Classification](https://savrn.com/datasets/tasks/text-classification) 38
- [Other](https://savrn.com/datasets/tasks/other) 34
- [Text Retrieval](https://savrn.com/datasets/tasks/text-retrieval) 33
- [Question Answering](https://savrn.com/datasets/tasks/question-answering) 31
- [Speech Recognition](https://savrn.com/datasets/tasks/speech-recognition) 23
- [Time Series Forecasting](https://savrn.com/datasets/tasks/time-series-forecasting) 18
- [Image Classification](https://savrn.com/datasets/tasks/image-classification) 15
- [Image to Text](https://savrn.com/datasets/tasks/image-to-text) 14
- [Tabular Classification](https://savrn.com/datasets/tasks/tabular-classification) 12
- [Visual Question Answering](https://savrn.com/datasets/tasks/visual-question-answering) 12

[See all](https://savrn.com/datasets/tasks)
