Home › Datasets › modality › Audio

Audio datasets

archive 2025-07-28

480 datasets carry the modality tag "Audio", ordered by the archive's paper count. Page 2 of 10: 48 shown of 480. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Audio datasets 49–96 of 480

VOICES (Voices Obscured In Complex Environmental Settings)
The VOICES corpus is a dataset to promote speech and signal processing research of speech recorded by far-field microphones in noisy room conditions.
48 papers · 0 benchmarks
MedleyDB, is a dataset of annotated, royalty-free multitrack recordings.
47 papers · 0 benchmarks
AliMeeting (Multi-Channel Multi-Party Meeting Transcription Challenge)
AliMeeting corpus consists of 120 hours of recorded Mandarin meeting data, including far-field data collected by 8-channel microphone array as well as near-field data collected by headset microphone.
46 papers · 1 benchmark
For understanding multimodal language used in expressing humor.
44 papers · 0 benchmarks
POP909 is a dataset which contains multiple versions of the piano arrangements of 909 popular songs created by professional musicians.
43 papers · 0 benchmarks
RWC (Real World Computing Music Database)
The RWC (Real World Computing) Music Database is a copyright-cleared music database (DB) that is available to researchers as a common foundation for research.
43 papers · 0 benchmarks
AVSpeech is a large-scale audio-visual dataset comprising speech clips with no interfering background signals.
42 papers · 0 benchmarks
The UCR Time Series Archive - introduced in 2002, has become an important resource in the time series data mining community, with at least one thousand published papers making use of at least one data set from the archive.
42 papers · 2 benchmarks
A ChatGPT-Assisted Weakly-Labelled Audio Captioning Dataset for Audio-Language Multimodal Research.
41 papers · 0 benchmarks
Sound Dataset for Malfunctioning Industrial Machine Investigation and Inspection (MIMII) is a sound dataset of industrial machine sounds.
40 papers · 0 benchmarks
The MTG-Jamendo dataset is an open dataset for music auto-tagging.
39 papers · 0 benchmarks
Slakh2100 (Synthesized Lakh Dataset)
The Synthesized Lakh (Slakh) Dataset is a dataset for audio source separation that is synthesized from the Lakh MIDI Dataset v0.1 using professional-grade sample-based virtual instruments.
38 papers · 3 benchmarks
URMP (University of Rochester Multi-Modal Musical Performance)
URMP (University of Rochester Multi-Modal Musical Performance) is a dataset for facilitating audio-visual analysis of musical performances.
38 papers · 2 benchmarks
WaveFake is a dataset for audio deepfake detection.
38 papers · 0 benchmarks
The Lakh MIDI dataset is a collection of 176,581 unique MIDI files, 45,129 of which have been matched and aligned to entries in the Million Song Dataset.
36 papers · 0 benchmarks
CoVoST is a large-scale multilingual speech-to-text translation corpus.
34 papers · 0 benchmarks
2000 HUB5 English Evaluation Transcripts was developed by the Linguistic Data Consortium (LDC) and consists of transcripts of 40 English telephone conversations used in the 2000 HUB5 evaluation sponsored by NIST (National Institute of…
33 papers · 2 benchmarks
A new multitask action quality assessment (AQA) dataset, the largest to date, comprising of more than 1600 diving samples; contains detailed annotations for fine-grained action recognition, commentary generation, and estimating the AQA…
33 papers · 2 benchmarks
Music21 is an untrimmed video dataset crawled by keyword query from Youtube.
33 papers · 0 benchmarks
DCASE 2016 is a dataset for sound event detection.
31 papers · 0 benchmarks
GuitarSet is a dataset of high-quality guitar recordings and rich annotations.
31 papers · 2 benchmarks
VocalSet (VocalSet: A Singing Voice Dataset)
VocalSet is a a singing voice dataset consisting of 10.1 hours of monophonic recorded audio of professional singers demonstrating both standard and extended vocal techniques on all 5 vowels.
30 papers · 2 benchmarks
The COUGHVID dataset provides over 20,000 crowdsourced cough recordings representing a wide range of subject ages, genders, geographic locations, and COVID-19 statuses.
29 papers · 0 benchmarks
CREMA-D is an emotional multimodal actor data set of 7,442 original clips from 91 actors.
28 papers · 7 benchmarks
MuseData is an electronic library of orchestral and piano classical music from CCARH.
28 papers · 0 benchmarks
RAVDESS (Ryerson Audio-Visual Database of Emotional Speech and Song)
The Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS) contains 7,356 files (total size: 24.8 GB).
27 papers · 6 benchmarks
CVSS is a massively multilingual-to-English speech to speech translation (S2ST) corpus, covering sentence-level parallel S2ST pairs from 21 languages into English.
26 papers · 1 benchmark
VocalSound is a free dataset consisting of 21,024 crowdsourced recordings of laughter, sighs, coughs, throat clearing, sneezes, and sniffs from 3,365 unique subjects.
24 papers · 1 benchmark
BEAT2 (BEAT-SMPLX-FLAME)
We propose EMAGE, a framework to generate full-body human gestures from audio and masked gestures, encompassing facial, local body, hands, and global movements.
23 papers · 2 benchmarks
EMOPIA (A Multi-Modal Pop Piano Dataset For Emotion Recognition and Emotion-based Music Generation)
EMOPIA (pronounced ‘yee-mò-pi-uh’) dataset is a shared multi-modal (audio and MIDI) database focusing on perceived emotion in pop piano music, to facilitate research on various tasks related to music emotion.
23 papers · 0 benchmarks
The Easy Communications (EasyCom) dataset is a world-first dataset designed to help mitigate the cocktail party effect from an augmented-reality (AR) -motivated multi-sensor egocentric world view.
22 papers · 4 benchmarks
The LOCATA dataset is a dataset for acoustic source localization.
22 papers · 0 benchmarks
CAL500 (Computer Audition Lab 500)
CAL500 (Computer Audition Lab 500) is a dataset aimed for evaluation of music information retrieval systems.
21 papers · 0 benchmarks
ICBHI Respiratory Sound Database (The Respiratory Sound database - ICBHI 2017 Challenge)
The Respiratory Sound database was originally compiled to support the scientific challenge organized at Int.
21 papers · 2 benchmarks
MIR-1K (Multimedia Information Retrieval lab, 1000 song clips) is a dataset designed for singing voice separation.
21 papers · 0 benchmarks
STARSS23 (STARSS23: An Audio-Visual Dataset of Spatial Recordings of Real Scenes with Spatiotemporal Annotations of Sound Events)
The Sony-TAu Realistic Spatial Soundscapes 2023 (STARSS23) dataset contains multichannel recordings of sound scenes in various rooms and environments, together with temporal and spatial annotations of prominent events belonging to a set of…
21 papers · 0 benchmarks
Consists of 1106 action samples from seven actions with quality scores as measured by expert human judges.
20 papers · 1 benchmark
AVSBench (Audio −Visual Segmentation)
AVSBench is a pixel-level audio-visual segmentation benchmark that provides ground truth labels for sounding objects.
20 papers · 0 benchmarks
The DiCOVA Challenge dataset is derived from the Coswara dataset, a crowd-sourced dataset of sound recordings from COVID-19 positive and non-COVID-19 individuals.
20 papers · 1 benchmark
The iKala dataset is a singing voice separation dataset that comprises of 252 30-second excerpts sampled from 206 iKala songs (plus 100 hidden excerpts reserved for MIREX data mining contest).
20 papers · 1 benchmark
DIRHA (Distant-speech Interaction for Robust Home Applications)
DIRHA-English is a multi-microphone database composed of real and simulated sequences of 1-minute.
19 papers · 1 benchmark
The FSDnoisy18k dataset is an open dataset containing 42.5 hours of audio across 20 sound event classes, including a small amount of manually-labeled data and a larger quantity of real-world noisy data.
19 papers · 0 benchmarks
STARSS22 (Sony-TAu Realistic Spatial Soundscapes 2022)
The Sony-TAu Realistic Spatial Soundscapes 2022(STARSS22) dataset consists of recordings of real scenes captured with high channel-count spherical microphone array (SMA).
19 papers · 1 benchmark
DESED (Domestic environment sound event detection)
The DESED dataset is a dataset designed to recognize sound event classes in domestic environments.
17 papers · 1 benchmark
CMU-MOSI (Multimodal Corpus of Sentiment Intensity)
The Multimodal Corpus of Sentiment Intensity (CMU-MOSI) dataset is a collection of 2199 opinion video clips.
16 papers · 2 benchmarks
Casual Conversations dataset is designed to help researchers evaluate their computer vision and audio models for accuracy across a diverse set of age, genders, apparent skin tones and ambient lighting conditions.
16 papers · 0 benchmarks
DiPCo (DiPCo -- Dinner Party Corpus)
We present a speech data corpus that simulates a "dinner party" scenario taking place in an everyday home environment.
16 papers · 0 benchmarks
FUSS (Free Universal Sound Separation)
The Free Universal Sound Separation (FUSS) dataset is a database of arbitrary sound mixtures and source-level references, for use in experiments on arbitrary sound separation.
16 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.