Home › Datasets › modality › Audio

Audio datasets

archive 2025-07-28

480 datasets carry the modality tag "Audio", ordered by the archive's paper count. Page 5 of 10: 48 shown of 480. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Audio datasets 193–240 of 480

FINO-Net is a multimodal (RGB, depth and audio) dataset, containing 229 real-world manipulation data of 5 different manipulation types recorded with a Baxter robot.
4 papers · 0 benchmarks
The FakeMusicCaps dataset contains total of 27605 10 seconds music tracks corresponding to almost 77 hours, generated using 5 different Text-To-Music (TTM) models.
4 papers · 0 benchmarks
GoodSounds dataset contains around 28 hours of recordings of single notes and scales played by 15 different professional musicians, all of them holding a music degree and having some expertise in teaching.
4 papers · 0 benchmarks
Hate speech has become one of the most significant issues in modern society, with implications in both the online and offline worlds.
4 papers · 1 benchmark
Human-Animal-Cartoon (HAC) dataset consists of seven actions (‘sleeping’, ‘watching tv’, ‘eating’, ‘drinking’, ‘swimming’, ‘running’, and ‘opening door’) performed by humans, animals, and cartoon figures, forming three different domains.
4 papers · 0 benchmarks
ITALIC: An ITALian Intent Classification Dataset ITALIC is an intent classification dataset for the Italian language, which is the first of its kind.
4 papers · 0 benchmarks
The M5Product dataset is a large-scale multi-modal pre-training dataset with coarse and fine-grained annotations for E-products.
4 papers · 0 benchmarks
MTASS is an open-source dataset in which mixtures contain three types of audio signals.
4 papers · 0 benchmarks
MuMu is a new dataset of more than 31k albums classified into 250 genre classes.
4 papers · 0 benchmarks
For each dataset we provide a short description as well as some characterization metrics.
4 papers · 0 benchmarks
NIPS4Bplus is a richly annotated birdsong audio dataset, that is comprised of recordings containing bird vocalisations along with their active species tags plus the temporal annotations acquired for them.
4 papers · 0 benchmarks
The Nottingham Dataset is a collection of 1200 American and British folk songs.
4 papers · 1 benchmark
The OLGA dataset contains artist similarities from AllMusic, together with content features from AcousticBrainz.
4 papers · 0 benchmarks
A dataset for urban sound tagging with spatiotemporal information.
4 papers · 0 benchmarks
TAU Spatial Sound Events 2019 consists of 2 datasets: Ambisonic (FOA) and Microphone Array (MIC), of identical sound scenes with the only difference in the format of the audio.
4 papers · 0 benchmarks
The TAU-NIGENS Spatial Sound Events 2021 dataset contains multiple spatial sound-scene recordings, consisting of sound events of distinct categories integrated into a variety of acoustical spaces, and from multiple source directions and…
4 papers · 1 benchmark
VietMed (VietMed: A Dataset and Benchmark for Automatic Speech Recognition of Vietnamese in the Medical Domain)
We introduced a Vietnamese speech recognition dataset in the medical domain comprising 16h of labeled medical speech, 1000h of unlabeled medical speech and 1200h of unlabeled general-domain speech.
4 papers · 2 benchmarks
Warblr is a dataset for the acoustic detection of birds.
4 papers · 0 benchmarks
dMelodies is dataset of simple 2-bar melodies generated using 9 independent latent factors of variation where each data point represents a unique melody based on the following constraints: - Each melody will correspond to a unique scale…
4 papers · 0 benchmarks
jazznet is a dataset of piano patterns for music audio machine learning research.
4 papers · 0 benchmarks
Due to the highly variable sample size of the original BirdClef2020 dataset and the issues that it presents with reproducibility, we propose a pruned version of the set, where samples longer than 180s are removed along with classes with…
3 papers · 1 benchmark
The BirdVox-full-night dataset contains 6 audio recordings, each about ten hours in duration.
3 papers · 0 benchmarks
COSIAN (a collection of singing voice annotation)
COSIAN is an annotation collection of Japanese popular (J-POP) songs, focusing on singing style and expression of famous solo-singers.
3 papers · 0 benchmarks
DCASE2014 is an audio classification benchmark.
3 papers · 0 benchmarks
E-GMD (Expanded Groove MIDI Dataset)
Expanded Groove MIDI dataset (E-GMD) is an automatic drum transcription (ADT) dataset that contains 444 hours of audio from 43 drum kits, making it an order of magnitude larger than similar datasets, and the first with human-performed…
3 papers · 0 benchmarks
ENST Drums (ENST-Drums: an extensive audio-visual database for drum signals processing)
ENST-Drums: an extensive audio-visual database for drum signals processing Olivier Gillet and Gaël Richard GET / ENST, CNRS LTCI, 37 rue Dareau, 75014 Paris, France The ENST-Drums database is a large and varied research database for…
3 papers · 0 benchmarks
FSDD (Free Spoken Digit Dataset)
Free Spoken Digit Dataset (FSDD) is a simple audio/speech dataset consisting of recordings of spoken digits in wav files at 8kHz.
3 papers · 0 benchmarks
Fraxtil is an audio dataset where given a raw audio track, the goal is to produce a choreography step chart, similar to those used in the Dance Dance Revolution video game.
3 papers · 0 benchmarks
The dataset consists of the features associated with 402 5-second sound samples.
3 papers · 0 benchmarks
HUI speech corpus (Hof University iisys speech dataset)
The data set contains several speakers.
3 papers · 2 benchmarks
ITG (In The Groove)
In The Groove (ITG) is an audio dataset where given a raw audio track, the goal is to produce a choreography step chart, similar to those used in the Dance Dance Revolution video game.
3 papers · 0 benchmarks
InfantMarmosetsVox is a dataset for multi-class call-type and caller identification.
3 papers · 0 benchmarks
The Jamendo Corpus is a voice detection dataset consisting of 93 songs with Creative Commons license from the Jamendo free music sharing website.
3 papers · 0 benchmarks
The LITIS-Rouen dataset is a dataset for audio scenes.
3 papers · 0 benchmarks
MISP2021 (Multimodal Information Based Speech Processing 2021)
The MISP2021 challenge dataset is a collection of audio-visual conversational data recorded in a home TV scenario using distant multi-microphones.
3 papers · 0 benchmarks
Operating rooms (ORs) are complex, high-stakes environments requiring precise understanding of interactions among medical staff, tools, and equipment for enhancing surgical assistance, situational awareness, and patient safety.
3 papers · 3 benchmarks
MMDB (Multimodal Dyadic Behavior)
Multimodal Dyadic Behavior (MMDB) dataset is a unique collection of multimodal (video, audio, and physiological) recordings of the social and communicative behavior of toddlers.
3 papers · 0 benchmarks
New refined labels for the MusicNet dataset obtained by the EM process as described in the paper: Ben Maman and Amit Bermano, "Unaligned Supervision for Automatic Music Transcription in The Wild"
3 papers · 0 benchmarks
ODSQA (Open-Domain Spoken Question Answering)
The ODSQA dataset is a spoken dataset for question answering in Chinese.
3 papers · 0 benchmarks
PACS (Physical Audiovisual CommonSense) is the first audiovisual benchmark annotated for physical commonsense attributes.
3 papers · 1 benchmark
RESD (Russian Emotional Speech Dialogs with annotated text)
Russian dataset of emotional speech dialogues.
3 papers · 1 benchmark
RWCP-SSD-Onomatopoeia is a dataset consisting of 155,568 onomatopoeic words paired with audio samples for environmental sound synthesis.
3 papers · 0 benchmarks
RemFX (RemFX Evaluation Datasets)
Audio samples processed with sound effects, to evaluate effect removal models.
3 papers · 0 benchmarks
SAVEE (Surrey Audio-Visual Expressed Emotion)
The Surrey Audio-Visual Expressed Emotion (SAVEE) dataset was recorded as a pre-requisite for the development of an automatic emotion recognition system.
3 papers · 1 benchmark
SONICS (Synthetic Or Not - Identifying Counterfeit Songs)
SONICS is a large-scale dataset comprising 97,164 songs — 48,090 real songs from YouTube and 49,074 fake songs from Suno & Udio — designed for synthetic song detection (SSD), also known as fake song detection (FSD).
3 papers · 0 benchmarks
A short clip of video may contain progression of multiple events and an interesting story line.
3 papers · 3 benchmarks
LA-2A Compressor data to accompany the paper "SignalTrain: Profiling Audio Compressors with Deep Neural Networks," https://arxiv.org/abs/1905.11928 Accompanying computer code: https://github.com/drscotthawley/signaltrain A collection of…
3 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.