Home › Datasets › task › Audio Classification
Audio Classification datasets
archive 2025-07-28
43 datasets carry the task tag "Audio Classification" (the task itself: Audio Classification), ordered by the archive's paper count. Page 1 of 1: 43 shown of 43. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Audio Classification datasets 1–43 of 43
The MNIST database (Modified National Institute of Standards and Technology database) is a large collection of handwritten digits.
7,651 papers · 44 benchmarks
Audioset is an audio event dataset, which consists of over 2M human-annotated 10-second video clips.
744 papers · 5 benchmarks
Common Voice is an audio dataset that consists of a unique MP3 and corresponding text file.
449 papers · 143 benchmarks
Speech Commands is an audio dataset of spoken words designed to help train and evaluate keyword spotting systems .
392 papers · 4 benchmarks
The ESC-50 dataset is a labeled collection of 2000 environmental audio recordings suitable for benchmarking methods of environmental sound classification.
387 papers · 4 benchmarks
Consists of more than 210k videos for 310 audio classes.
211 papers · 3 benchmarks
This paper introduces the pipeline to scale the largest dataset in egocentric vision EPIC-KITCHENS.
162 papers · 6 benchmarks
FSD50K (Freesound Database 50K)
Freesound Dataset 50k (or FSD50K for short) is an open dataset of human-labeled sound events containing 51,197 Freesound clips unequally distributed in 200 classes drawn from the AudioSet Ontology.
155 papers · 2 benchmarks
Urban Sound 8K is an audio dataset that contains 8732 labeled sound excerpts (<=4s) of urban sounds from 10 classes: airconditioner, carhorn, childrenplaying, dogbark, drilling, engingeidling, gunshot, jackhammer, siren, and streetmusic.
147 papers · 1 benchmark
The UCR Time Series Archive - introduced in 2002, has become an important resource in the time series data mining community, with at least one thousand published papers making use of at least one data set from the archive.
42 papers · 2 benchmarks
CREMA-D is an emotional multimodal actor data set of 7,442 original clips from 91 actors.
28 papers · 7 benchmarks
RAVDESS (Ryerson Audio-Visual Database of Emotional Speech and Song)
The Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS) contains 7,356 files (total size: 24.8 GB).
27 papers · 6 benchmarks
VocalSound is a free dataset consisting of 21,024 crowdsourced recordings of laughter, sighs, coughs, throat clearing, sneezes, and sniffs from 3,365 unique subjects.
24 papers · 1 benchmark
The Respiratory Sound database was originally compiled to support the scientific challenge organized at Int.
21 papers · 2 benchmarks
The DiCOVA Challenge dataset is derived from the Coswara dataset, a crowd-sourced dataset of sound recordings from COVID-19 positive and non-COVID-19 individuals.
20 papers · 1 benchmark
The FSDnoisy18k dataset is an open dataset containing 42.5 hours of audio across 20 sound event classes, including a small amount of manually-labeled data and a larger quantity of real-world noisy data.
19 papers · 0 benchmarks
SHD (Spiking Heidelberg Digits)
The Spiking Heidelberg Digits (SHD) dataset is an audio-based classification dataset of 1k spoken digits ranging from zero to nine in the English and German languages.
19 papers · 1 benchmark
The gtzan8 audio dataset contains 1000 tracks of 30 second length.
15 papers · 4 benchmarks
EPIC-SOUNDS is a large scale dataset of audio annotations capturing temporal extents and class labels within the audio stream of the egocentric videos from EPIC-KITCHENS-100.
12 papers · 2 benchmarks
The YouTube-100M data set consists of 100 million YouTube videos: 70M training videos, 10M evaluation videos, and 20M validation videos.
8 papers · 0 benchmarks
SSC (Spiking Speech Commands v0.2)
The SSC dataset is a spiking version of the Speech Commands dataset release by Google (Speech Commands).
7 papers · 1 benchmark
The dataset uses VGG-Sound which consists of 10s clips collected from YouTube for 309 sound classes.
5 papers · 0 benchmarks
The aGender corpus contains audio recordings of predefined utterances and free speech produced by humans of different age and gender.
5 papers · 0 benchmarks
A dataset for urban sound tagging with spatiotemporal information.
4 papers · 0 benchmarks
The TAU-NIGENS Spatial Sound Events 2021 dataset contains multiple spatial sound-scene recordings, consisting of sound events of distinct categories integrated into a variety of acoustical spaces, and from multiple source directions and…
4 papers · 1 benchmark
DCASE2014 is an audio classification benchmark.
3 papers · 0 benchmarks
HUME-VB (The Hume Vocal Bursts Dataset)
The Hume Vocal Burst Database (H-VB) includes all train, validation, and test recordings and corresponding emotion ratings for the train and validation recordings.
3 papers · 7 benchmarks
InfantMarmosetsVox is a dataset for multi-class call-type and caller identification.
3 papers · 0 benchmarks
Annotated audio files (separate combined annotation file) of lung sounds as recorded from various vantage points of the chest wall.
2 papers · 1 benchmark
This repository contains the SINGA:PURA dataset, a strongly-labelled polyphonic urban sound dataset with spatiotemporal context.
2 papers · 0 benchmarks
Overview nEMO is a simulated dataset of emotional speech in the Polish language.
2 papers · 0 benchmarks
We recorded gun sounds by changing the type and position of guns to diversify distances and angles in the PUBG environment.
1 paper · 0 benchmarks
DEEP-VOICE: Real-time Detection of AI-Generated Speech for DeepFake Voice Conversion This dataset contains examples of real human speech, and DeepFake versions of those speeches by using Retrieval-based Voice Conversion.
1 paper · 1 benchmark
MediBeng (Synthetic Code-Switched Bengali-English Speech Conversations for Healthcare Applications)
MediBeng Dataset The MediBeng dataset contains synthetic code-switched dialogues in Bengali and English for training models in speech recognition (ASR), text-to-speech (TTS), and machine translation in clinical settings.
1 paper · 1 benchmark
A large-scale reference dataset for bioacoustics.
1 paper · 1 benchmark
Dataset for multimodal skills assessment focusing on assessing piano player’s skill level.
1 paper · 3 benchmarks
PC-GITA is a Spanish speech corpus designed to analyze speech impairments in individuals with Parkinson's Disease (PD).
1 paper · 0 benchmarks
Asthma is a common, usually long-term respiratory disease with negative impact on society and the economy worldwide.
1 paper · 0 benchmarks
The full version of ReefSet used in Williams et al.
1 paper · 0 benchmarks
arxiv : https://arxiv.org/abs/2304.11708 Accepted at 29th International Congress on Sound and Vibration (ICSV29).
1 paper · 1 benchmark
The Humbug Zooinverse dataset is a dataset of mosquito audio recordings.
1 paper · 0 benchmarks
Dataset Summary The Deep Evaluation of Audio Representations (DEAR) dataset is a benchmark designed to assess general-purpose audio foundation models on properties critical for hearable devices.
0 papers · 0 benchmarks
Mudestreda (Mudestreda Multimodal Device State Recognition Dataset)
Mudestreda Multimodal Device State Recognition Dataset obtained from real industrial milling device with Time Series and Image Data for Classification, Regression, Anomaly Detection, Remaining Useful Life (RUL) estimation, Signal Drift…
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.