Home › Datasets › modality › Audio

Audio datasets

archive 2025-07-28

480 datasets carry the modality tag "Audio", ordered by the archive's paper count. Page 1 of 10: 48 shown of 480. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Audio datasets 1–48 of 480

The LibriSpeech corpus is a collection of approximately 1,000 hours of audiobooks that are a part of the LibriVox project.
2,361 papers · 4 benchmarks
IEMOCAP (The Interactive Emotional Dyadic Motion Capture (IEMOCAP) Database)
Multimodal Emotion Recognition IEMOCAP The IEMOCAP dataset consists of 151 videos of recorded dialogues, with 2 speakers per session for a total of 302 videos across the dataset.
749 papers · 3 benchmarks
Audioset is an audio event dataset, which consists of over 2M human-annotated 10-second video clips.
744 papers · 5 benchmarks
VoxCeleb1 is an audio dataset containing over 100,000 utterances for 1,251 celebrities, extracted from videos uploaded to YouTube.
680 papers · 10 benchmarks
VoxCeleb2 is a large scale speaker recognition dataset obtained automatically from open-source media.
564 papers · 5 benchmarks
The Universal Dependencies (UD) project seeks to develop cross-linguistically consistent treebank annotation of morphology and syntax for multiple languages.
520 papers · 5 benchmarks
VCTK (CSTR VCTK Corpus)
This CSTR VCTK Corpus includes speech data uttered by 110 English speakers with various accents.
476 papers · 6 benchmarks
Common Voice is an audio dataset that consists of a unique MP3 and corresponding text file.
449 papers · 143 benchmarks
Speech Commands is an audio dataset of spoken words designed to help train and evaluate keyword spotting systems .
392 papers · 4 benchmarks
The ESC-50 dataset is a labeled collection of 2000 environmental audio recordings suitable for benchmarking methods of environmental sound classification.
387 papers · 4 benchmarks
LJSpeech (The LJ Speech Dataset)
This is a public domain speech dataset consisting of 13,100 short audio clips of a single speaker reading passages from 7 non-fiction books.
323 papers · 2 benchmarks
AudioCaps is a dataset of sounds with event descriptions that was introduced for the task of audio captioning, with sounds sourced from the AudioSet dataset.
279 papers · 6 benchmarks
LibriTTS is a multi-speaker English corpus of approximately 585 hours of read English speech at 24kHz sampling rate, prepared by Heiga Zen with the assistance of Google Speech and Google Brain team members.
257 papers · 1 benchmark
MuST-C currently represents the largest publicly available multilingual corpus (one-to-many) for speech translation.
216 papers · 2 benchmarks
MUSAN is a corpus of music, speech and noise.
204 papers · 0 benchmarks
Clotho is an audio captioning dataset, consisting of 4981 audio samples, and each audio sample has five captions (a total of 24 905 captions).
202 papers · 3 benchmarks
CMU Multimodal Opinion Sentiment and Emotion Intensity (CMU-MOSEI) is the largest dataset of sentence-level sentiment analysis and emotion recognition in online videos.
190 papers · 3 benchmarks
LRW (Lip Reading in the Wild)
The Lip Reading in the Wild (LRW) dataset a large-scale audio-visual database that contains 500 different words from over 1,000 speakers.
188 papers · 8 benchmarks
FSD50K (Freesound Database 50K)
Freesound Dataset 50k (or FSD50K for short) is an open dataset of human-labeled sound events containing 51,197 Freesound clips unequally distributed in 200 classes drawn from the AudioSet Ontology.
155 papers · 2 benchmarks
Urban Sound 8K is an audio dataset that contains 8732 labeled sound excerpts (<=4s) of urban sounds from 10 classes: airconditioner, carhorn, childrenplaying, dogbark, drilling, engingeidling, gunshot, jackhammer, siren, and streetmusic.
147 papers · 1 benchmark
The SumMe dataset is a video summarization dataset consisting of 25 videos, each annotated with at least 15 human summaries (390 in total).
146 papers · 3 benchmarks
FLEURS (Few-shot Learning Evaluation of Universal Representations of Speech)
We introduce FLEURS, the Few-shot Learning Evaluation of Universal Representations of Speech benchmark.
141 papers · 1 benchmark
NSynth is a dataset of one shot instrumental notes, containing 305,979 musical notes with unique pitch, timbre and envelope.
138 papers · 2 benchmarks
FMA (Free Music Archive)
The Free Music Archive (FMA) is a large-scale dataset for evaluating several tasks in Music Information Retrieval.
128 papers · 2 benchmarks
LSMDC (Large Scale Movie Description Challenge)
This dataset contains 118,081 short video clips extracted from 202 movies.
126 papers · 3 benchmarks
The MAESTRO dataset contains over 200 hours of paired audio and MIDI recordings from ten years of International Piano-e-Competition.
118 papers · 1 benchmark
LRS2 (Lip Reading Sentences 2)
The Oxford-BBC Lip Reading Sentences 2 (LRS2) dataset is one of the largest publicly available datasets for lip reading sentences in-the-wild.
115 papers · 10 benchmarks
The MUSDB18 is a dataset of 150 full lengths music tracks (~10h duration) of different genres along with their isolated drums, bass, vocals and others stems.
106 papers · 2 benchmarks
Sleep-EDF (Sleep-EDF Expanded)
The sleep-edf database contains 197 whole-night PolySomnoGraphic sleep recordings, containing EEG, EOG, chin EMG, and event markers.
94 papers · 5 benchmarks
GigaSpeech, an evolving, multi-domain English speech recognition corpus with 10,000 hours of high quality labeled audio suitable for supervised training, and 40,000 hours of total audio suitable for semi-supervised and unsupervised…
87 papers · 3 benchmarks
The How2 dataset contains 13,500 videos, or 300 hours of speech, and is split into 185,187 training, 2022 development (dev), and 2361 test utterances.
84 papers · 2 benchmarks
MetaQA (MoviE Text Audio QA)
The MetaQA dataset consists of a movie ontology derived from the WikiMovies Dataset and three sets of question-answer pairs written in natural language: 1-hop, 2-hop, and 3-hop queries.
81 papers · 1 benchmark
VQG (Visual Question Generation)
VQG is a collection of datasets for visual question generation.
80 papers · 1 benchmark
Multilingual LibriSpeech is a large multilingual corpus suitable for speech research.
77 papers · 2 benchmarks
MSP-IMPROV (MSP-IMPROV: An Acted Corpus of Dyadic Interactions to Study Emotion Perception)
We present the MSP-IMPROV corpus, a multimodal emotional database, where the goal is to have control over lexical content and emotion while also promoting naturalness in the recordings.
70 papers · 1 benchmark
MagnaTagATune dataset contains 25,863 music clips.
65 papers · 2 benchmarks
The TED-LIUM corpus consists of English-language TED talks.
64 papers · 2 benchmarks
We propose Localized Narratives, a new form of multimodal image annotations connecting vision and language.
63 papers · 3 benchmarks
EmoryNLP comprises 97 episodes, 897 scenes, and 12,606 utterances, where each utterance is annotated with one of the seven emotions borrowed from the six primary emotions in the Willcox (1982)’s feeling wheel, sad, mad, scared, powerful,…
58 papers · 1 benchmark
XD-Violence is a large-scale audio-visual dataset for violence detection in videos.
58 papers · 2 benchmarks
Fluent Speech Commands is an open source audio dataset for spoken language understanding (SLU) experiments.
57 papers · 1 benchmark
WHAMR! (WHAM! with synthetic reverberated sources)
WHAMR!
57 papers · 3 benchmarks
BEAT (Body-Expression-Audio-Text)
BEAT has i) 76 hours, high-quality, multi-modal data captured from 30 speakers talking with eight different emotions and in four different languages, ii) 32 millions frame-level emotion and semantic relevance annotations.
56 papers · 1 benchmark
ReVerb Challenge (REverberant Voice Enhancement and Recognition Benchmark)
The REVERB (REverberant Voice Enhancement and Recognition Benchmark) challenge is a benchmark for evaluation of automatic speech recognition techniques.
55 papers · 1 benchmark
The SEMAINE videos dataset contains spontaneous data capturing the audiovisual interaction between a human and an operator undertaking the role of an avatar with four personalities: Poppy (happy), Obadiah (gloomy), Spike (angry) and…
54 papers · 1 benchmark
VoiceBank + DEMAND (Noisy speech database for training speech enhancement algorithms and TTS models)
VoiceBank+DEMAND is a noisy speech database for training speech enhancement algorithms and TTS models.
53 papers · 1 benchmark
The large-scale MUSIC-AVQA dataset of musical performance contains 45,867 question-answer pairs, distributed in 9,288 videos for over 150 hours.
51 papers · 1 benchmark

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.