Home › Datasets › modality › Audio
Audio datasets
archive 2025-07-28
480 datasets carry the modality tag "Audio", ordered by the archive's paper count. Page 4 of 10: 48 shown of 480. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Audio datasets 145–192 of 480
CochlScene is a dataset for acoustic scene classification.
7 papers · 1 benchmark
The CocoChorales Dataset CocoChorales is a dataset consisting of over 1400 hours of audio mixtures containing four-part chorales performed by 13 instruments, all synthesized with realistic-sounding generative models.
7 papers · 0 benchmarks
MSSD (Music Streaming Sessions Dataset)
The Spotify Music Streaming Sessions Dataset (MSSD) consists of 160 million streaming sessions with associated user interactions, audio features and metadata describing the tracks streamed during the sessions, and snapshots of the…
7 papers · 1 benchmark
NES-MDB (Nintendo Entertainment System Music Database)
The Nintendo Entertainment System Music Database (NES-MDB) is a dataset intended for building automatic music composition systems for the NES audio synthesizer.
7 papers · 0 benchmarks
SSC (Spiking Speech Commands v0.2)
The SSC dataset is a spiking version of the Speech Commands dataset release by Google (Speech Commands).
7 papers · 1 benchmark
SingFake (SingFake: Singing Voice Deepfake Detection)
The rise of singing voice synthesis presents critical challenges to artists and industry stakeholders over unauthorized voice usage.
7 papers · 0 benchmarks
TUT-SED Synthetic 2016 contains of mixture signals artificially generated from isolated sound events samples.
7 papers · 0 benchmarks
ADVANCE (AuDio Visual Aerial sceNe reCognition datasEt)
The AuDio Visual Aerial sceNe reCognition datasEt (ADVANCE) is a brand-new multimodal learning dataset, which aims to explore the contribution of both audio and conventional visual messages to scene recognition.
6 papers · 0 benchmarks
ASCEND (A Spontaneous Chinese-English Dataset) introduces a high-quality resource of spontaneous multi-turn conversational dialogue Chinese code-switching corpus collected in Hong Kong.
6 papers · 0 benchmarks
CCMixter is a singing voice separation dataset consisting of 50 full-length stereo tracks from ccMixter featuring many different musical genres.
6 papers · 0 benchmarks
The CHB-MIT dataset is a dataset of EEG recordings from pediatric subjects with intractable seizures.
6 papers · 1 benchmark
The German Lipreading dataset consists of 250,000 publicly available videos of the faces of speakers of the Hessian Parliament, which was processed for word-level lip reading using an automatic pipeline.
6 papers · 0 benchmarks
GTSinger (GTSinger: A Global Multi-Technique Singing Corpus with Realistic Music Scores for All Singing Tasks)
The scarcity of high-quality and multi-task singing datasets significantly hinders the development of diverse controllable and personalized singing tasks, as existing singing datasets suffer from low quality, limited diversity of languages…
6 papers · 0 benchmarks
L3DAS21 is a dataset for 3D audio signal processing.
6 papers · 2 benchmarks
LSSED, a challenging large-scale english dataset for speech emotion recognition.
6 papers · 1 benchmark
The MEDIA French corpus is dedicated to semantic extraction from speech in a context of human/machine dialogues.
6 papers · 0 benchmarks
Moviescope is a large-scale dataset of 5,000 movies with corresponding video trailers, posters, plots and metadata.
6 papers · 0 benchmarks
MuChoMusic is a benchmark designed to evaluate music understanding in multimodal language models focused on audio.
6 papers · 0 benchmarks
The MusicBench dataset is a music audio-text pair dataset that was designed for text-to-music generation purpose and released along with Mustango text-to-music model.
6 papers · 1 benchmark
The ObjectFolder Real dataset contains multisensory data collected from 100 real-world household objects.
6 papers · 0 benchmarks
The data and audio included here were collected for the Soundscape Attributes Translation Project (SATP).
6 papers · 0 benchmarks
A set of approximately 100K podcast episodes comprised of raw audio files along with accompanying ASR transcripts.
6 papers · 0 benchmarks
WSJ0-2mix-extr is a speech extraction dataset
6 papers · 1 benchmark
Acappella comprises around 46 hours of a cappella solo singing videos sourced from YouTbe, sampled across different singers and languages.
5 papers · 0 benchmarks
AccentDB is a database that contains samples of 4 Indian-English accents, and a compilation of samples from 4 native-English, and a metropolitan Indian-English accent.
5 papers · 0 benchmarks
Artie Bias Corpus is an open dataset for detecting demographic bias in speech applications.
5 papers · 0 benchmarks
BRACE (The Breakdancing Competition Dataset for Dance Motion Synthesis)
BRACE is a dataset for audio-conditioned dance motion synthesis challenging common assumptions for this task: - strong music-dance correlation - controlled motion data - simple poses and movements To address these issues: - We focus on…
5 papers · 2 benchmarks
This data set includes beat and bar annotations of the ballroom dataset, introduced by Gouyon et al.
5 papers · 3 benchmarks
BIWI 3D corpus comprises a total of 1109 sentences uttered by 14 native English speakers (6 males and 8 females).
5 papers · 1 benchmark
The English data for voice building was obtained, prepared and provided the the challenge by Lessac Technologies Inc., having originally came from the publishers Voice Factory International Inc.
5 papers · 1 benchmark
CHiME-Home is a dataset for sound source recognition in a domestic environment.
5 papers · 0 benchmarks
TAU Urban Acoustic Scenes 2019 Mobile development dataset consists of 10-seconds audio segments from 10 acoustic scenes: Airport Indoor shopping mall Metro station Pedestrian street Public square Street with medium level of traffic…
5 papers · 1 benchmark
The Distress Analysis Interview Corpus/Wizard-of-Oz set (DAIC-WOZ) dataset [24, 25] comprises voice and text samples from 189 interviewed healthy and control persons and their PHQ-8 depression detection questionnaire.
5 papers · 0 benchmarks
ESC50 (ESC: Dataset for Environmental Sound Classification)
The ESC-50 dataset is a labeled collection of 2000 environmental audio recordings suitable for benchmarking methods of environmental sound classification.
5 papers · 0 benchmarks
FSDKaggle2019 is an audio dataset containing 29,266 audio files annotated with 80 labels of the AudioSet Ontology.
5 papers · 0 benchmarks
Dataset for lyrics alignment and transcription evaluation.
5 papers · 0 benchmarks
This dataset is a sound dataset for malfunctioning industrial machine investigation and inspection with domain shifts due to changes in operational and environmental conditions (MIMII DUE).
5 papers · 0 benchmarks
RealMAN (A Real-Recorded and Annotated Microphone Array Dataset for Dynamic Speech Enhancement and Localization)
The Audio Signal and Information Processing Lab at Westlake University, in collaboration with AISHELL, has released the Real-recorded and annotated Microphone Array speech&Noise (RealMAN) dataset, which provides annotated multi-channel…
5 papers · 2 benchmarks
The Song Describer Dataset (SDD) contains ~1.1k captions for 706 permissively licensed music recordings.
5 papers · 1 benchmark
ToyADMOS2 is a dataset of miniature-machine operating sounds for anomalous sound detection under domain shift conditions.
5 papers · 0 benchmarks
The dataset uses VGG-Sound which consists of 10s clips collected from YouTube for 309 sound classes.
5 papers · 0 benchmarks
The aGender corpus contains audio recordings of predefined utterances and free speech produced by humans of different age and gender.
5 papers · 0 benchmarks
ARAUS (Affective Responses to Augmented Urban Soundscapes)
Choosing optimal maskers for existing soundscapes to effect a desired perceptual change via soundscape augmentation is non-trivial due to extensive varieties of maskers and a dearth of benchmark datasets with which to compare and develop…
4 papers · 0 benchmarks
The Bach Doodle Dataset is composed of 21.6 million harmonizations submitted from the Bach Doodle.
4 papers · 0 benchmarks
This dataset includes the beat and downbeat annotations for Beatles albums.
4 papers · 2 benchmarks
Children's Song Dataset is open source dataset for singing voice research.
4 papers · 0 benchmarks
ComMU has 11,144 MIDI samples that consist of short note sequences created by professional composers with their corresponding 12 metadata.
4 papers · 0 benchmarks
DISCO-10M is a novel and extensive music dataset that surpasses the largest previously available music dataset by an order of magnitude.
4 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.