Home › Datasets › modality › Audio

Audio datasets

archive 2025-07-28

480 datasets carry the modality tag "Audio", ordered by the archive's paper count. Page 3 of 10: 48 shown of 480. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Audio datasets 97–144 of 480

Groove (Groove MIDI Dataset)
The Groove MIDI Dataset (GMD) is composed of 13.6 hours of aligned MIDI and (synthesized) audio of human-performed, tempo-aligned expressive drumming.
16 papers · 2 benchmarks
MeetingBank, a benchmark dataset created from the city councils of 6 major U.S.
16 papers · 1 benchmark
SONAR, a new multilingual and multimodal fixed-size sentence embedding space, with a full suite of speech and text encoders and decoders.
16 papers · 0 benchmarks
CPED (Chinese Personalized and Emotional Dialogue)
We construct a dataset named CPED from 40 Chinese TV shows.
15 papers · 3 benchmarks
The gtzan8 audio dataset contains 1000 tracks of 30 second length.
15 papers · 4 benchmarks
MUSDB18-HQ is a high-quality version of the MUSDB18 music tracks dataset.
15 papers · 1 benchmark
A large-scale dataset for retrieval and event localisation in video.
15 papers · 1 benchmark
ToyADMOS dataset is a machine operating sounds dataset of approximately 540 hours of normal machine operating sounds and over 12,000 samples of anomalous sounds collected with four microphones at a 48kHz sampling rate, prepared by Yuma…
15 papers · 0 benchmarks
AVSD (Audio-Visual Scene-Aware Dialog)
The Audio Visual Scene-Aware Dialog (AVSD) dataset, or DSTC7 Track 3, is a audio-visual dataset for dialogue understanding.
14 papers · 1 benchmark
DailyTalk is a high-quality conversational speech dataset designed for Text-to-Speech.
14 papers · 0 benchmarks
LAV-DF (Localized Audio Visual DeepFake Dataset)
Localized Audio Visual DeepFake Dataset (LAV-DF).
14 papers · 1 benchmark
TAU Urban Acoustic Scenes 2019 development dataset consists of 10-seconds audio segments from 10 acoustic scenes: airport, indoor shopping mall, metro station, pedestrian street, public square, street with medium level of traffic,…
14 papers · 2 benchmarks
ASAP (Aligned Scores and Performances)
ASAP is a dataset of 222 digital musical scores aligned with 1068 performances (more than 92 hours) of Western classical piano music.
13 papers · 2 benchmarks
L3DAS22: MACHINE LEARNING FOR 3D AUDIO SIGNAL PROCESSING This dataset supports the L3DAS22 IEEE ICASSP Gand Challenge.
13 papers · 0 benchmarks
The TUT Acoustic Scenes 2017 dataset is a collection of recordings from various acoustic scenes all from distinct locations.
13 papers · 1 benchmark
Existing audio-visual event localization (AVE) handles manually trimmed videos with only a single instance in each of them.
13 papers · 1 benchmark
EPIC-SOUNDS is a large scale dataset of audio annotations capturing temporal extents and class labels within the audio stream of the egocentric videos from EPIC-KITCHENS-100.
12 papers · 2 benchmarks
FSDKaggle2018 is an audio dataset containing 11,073 audio files annotated with 41 labels of the AudioSet Ontology.
12 papers · 1 benchmark
Neptune (Neptune Long Video Understanding Benchmark)
Neptune is a dataset consisting of challenging question-answer-decoy (QAD) sets for long videos (up to 15 minutes).
12 papers · 0 benchmarks
Co-speech gestures are everywhere.
12 papers · 1 benchmark
DCASE 2013 is a dataset for sound event detection.
11 papers · 0 benchmarks
the YF-E6 emotion dataset using the 6 basic emotion type as keywords on social video-sharing websites including YouTube and Flickr, leading to a total of 3000 videos.
11 papers · 1 benchmark
MSD (Million Song Dataset)
The Million Song Dataset is a freely-available collection of audio features and metadata for a million contemporary popular music tracks.
11 papers · 2 benchmarks
PATS (Pose Audio Transcript Style)
PATS dataset consists of a diverse and large amount of aligned pose, audio and transcripts.
11 papers · 0 benchmarks
Overall duration per microphone: about 36 hours (31 hrs train / 2.5 hrs dev / 2.5 hrs test) Count of microphones: 3 (Microsoft Kinect, Yamaha, Samson) Count of wave-files per microphone: about 14500 Overall count of participations: 180…
11 papers · 1 benchmark
VoxForge is an open speech dataset that was set up to collect transcribed speech for use with Free and Open Source Speech Recognition Engines (on Linux, Windows and Mac).
11 papers · 9 benchmarks
DCASE2018 Task 4 is a dataset for large-scale weakly labeled semi-supervised sound event detection in domestic environments.
10 papers · 0 benchmarks
The Lakh Pianoroll Dataset (LPD) is a collection of 174,154 multitrack pianorolls derived from the Lakh MIDI Dataset (LMD).
10 papers · 0 benchmarks
MACS (Multi-Annotator Captioned Soundscapes)
This is a dataset containing audio captions and corresponding audio tags for a number of 3930 audio files of the TAU Urban Acoustic Scenes 2019 development dataset (airport, public square, and park).
10 papers · 0 benchmarks
MUGEN is a large-scale video-audio-text dataset MUGEN, collected using the open-sourced platform game CoinRun.
10 papers · 0 benchmarks
MeshRIR is a dataset of acoustic room impulse responses (RIRs) at finely meshed grid points.
10 papers · 0 benchmarks
FAIR-Play is a video-audio dataset consisting of 1,871 video clips and their corresponding binaural audio clips recording in a music room.
9 papers · 0 benchmarks
Home Action Genome is a large-scale multi-view video database of indoor daily activities.
9 papers · 2 benchmarks
KSoF (The Kassel State of Fluency Dataset – A Therapy Centered Dataset of Stuttering)
Stuttering is a complex speech disorder that negatively affects an individual’s ability to communicate effectively.
9 papers · 0 benchmarks
MSP-Podcast (A large naturalistic speech emotional dataset)
The MSP-Podcast corpus contains speech segments from podcast recordings which are perceptually annotated using crowdsourcing.
9 papers · 4 benchmarks
MuSe-CaR (Multimodal Sentiment Analysis in Car Reviews)
The MuSe-CAR database is a large, multimodal (video, audio, and text) dataset which has been gathered in-the-wild with the intention of further understanding Multimodal Sentiment Analysis in-the-wild, e.g., the emotional engagement that…
9 papers · 0 benchmarks
The SALMon dataset and benchmark was introduced in the paper "A Suite for Acoustic Language Model Evaluation", with the goal of evaluating the modelling abilities of speech language models with regards to different kinds of acoustic…
9 papers · 1 benchmark
Speech encompasses a wealth of information, including but not limited to content, paralinguistic, and environmental information.
9 papers · 0 benchmarks
TUM-GAID (TUM Gait from Audio, Image and Depth) collects 305 subjects performing two walking trajectories in an indoor environment.
9 papers · 0 benchmarks
UDIVA is a new non-acted dataset of face-to-face dyadic interactions, where interlocutors perform competitive and collaborative tasks with different behavior elicitation and cognitive workload.
9 papers · 0 benchmarks
ADIMA is a novel, linguistically diverse, ethically sourced, expert annotated and well-balanced multilingual profanity detection audio dataset comprising of 11,775 audio samples in 10 Indic languages spanning 65 hours and spoken by 6,446…
8 papers · 0 benchmarks
OpenMIC-2018 is an instrument recognition dataset containing 20,000 examples of Creative Commons-licensed music available on the Free Music Archive.
8 papers · 1 benchmark
The SmartLights benchmark from Snipstests the capability of controlling lights in different rooms.
8 papers · 1 benchmark
We introduce a new audio dataset called SoundDescs that can be used for tasks such as text to audio retrieval, audio captioning etc.
8 papers · 1 benchmark
SoundingEarth consists of co-located aerial imagery and audio samples all around the world.
8 papers · 1 benchmark
The TUT Sound Events 2017 dataset contains 24 audio recordings in a street environment and contains 6 different classes.
8 papers · 0 benchmarks
The YouTube-100M data set consists of 100 million YouTube videos: 70M training videos, 10M evaluation videos, and 20M validation videos.
8 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.