Home › Datasets › modality › Audio
Audio datasets
archive 2025-07-28
480 datasets carry the modality tag "Audio", ordered by the archive's paper count. Page 6 of 10: 48 shown of 480. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Audio datasets 241–288 of 480
The SmartSpeaker benchmark tests the performance of reacting to music player commands in English as well as in French.
3 papers · 1 benchmark
The Tongue and Lips (TaL) corpus is a multi-speaker corpus of ultrasound images of the tongue and video images of lips.
3 papers · 0 benchmarks
Voice conversion (VC) is a technique to transform a speaker identity included in a source speech waveform into a different one while preserving linguistic information of the source speech waveform.
3 papers · 0 benchmarks
This work introduces Zambezi Voice, an open-source multilingual speech resource for Zambian languages.
3 papers · 0 benchmarks
AIME (AI Music Evaluation Dataset)
The AIME dataset contains 6,000 audio tracks generated by 12 music generation models in addition to 500 tracks from MTG-Jamendo.
2 papers · 0 benchmarks
ARCA23K is a dataset of labelled sound events created to investigate real-world label noise.
2 papers · 0 benchmarks
AV Digits Database is an audiovisual database which contains normal, whispered and silent speech.
2 papers · 0 benchmarks
AVCAffe (A Large Scale Audio-Visual Dataset of Cognitive Load and Affect for Remote Work)
We introduce AVCAffe, the first Audio-Visual dataset consisting of Cognitive load and Affect attributes.
2 papers · 0 benchmarks
AVSync15 is a high-quality synchronized audio-video dataset curated from VGGSound.
2 papers · 0 benchmarks
Audio-alpaca: A preference dataset for aligning text-to-audio models Audio-alpaca is a pairwise preference dataset containing about 15k (prompt,chosen, rejected) triplets where given a textual prompt, chosen is the preferred generated…
2 papers · 0 benchmarks
BiGe (Bielefeld Gesture Corpus)
The BiGe corpus is comprised of 54.360 shots of interest extracted from TED and TEDx talks.
2 papers · 0 benchmarks
BirdClef 2018 is a bird soundscape dataset based on the contributions of the Xeno-canto network.
2 papers · 0 benchmarks
BirdClef 2019 is a bird soundscape dataset.
2 papers · 0 benchmarks
Annotated audio files (separate combined annotation file) of lung sounds as recorded from various vantage points of the chest wall.
2 papers · 1 benchmark
The DCASE 2017 rare sound events dataset contains isolated sound events for three classes: 148 crying babies (mean duration 2.25s), 139 glasses breaking (mean duration 1.16s), and 187 gun shots (mean duration 1.32s).
2 papers · 0 benchmarks
Dusha (Dusha Crowd, Dusha Podcast)
Dusha is a dataset for speech emotion recognition (SER) tasks.
2 papers · 2 benchmarks
1000 songs has been selected from Free Music Archive (FMA).
2 papers · 1 benchmark
This dataset includes all music sources, background noises and impulse-reponses (IR) samples and conversation speech that have been used in the work "Neural Audio Fingerprint for High-specific Audio Retrieval based on Contrastive Learning"…
2 papers · 0 benchmarks
The Flickr 8k Audio Caption Corpus contains 40,000 spoken captions of 8,000 natural images.
2 papers · 0 benchmarks
JVS-MuSiC is a Japanese multispeaker singing-voice corpus called "JVS-MuSiC" with the aim to analyze and synthesize a variety of voices.
2 papers · 0 benchmarks
Jam-ALT (JamALT: A Formatting-Aware Lyrics Transcription Benchmark)
JamALT is a revision of the JamendoLyrics dataset (80 songs in 4 languages), adapted for use as an automatic lyrics transcription (ALT) benchmark.
2 papers · 5 benchmarks
This is a subset of Kinetics-400, introduced in Look, Listen and Learn by Relja Arandjelovic and Andrew Zisserman.
2 papers · 0 benchmarks
LUMA (Learning from Uncertain and Multimodal Data)
LUMA is a multimodal dataset that consists of audio, image, and text modalities.
2 papers · 0 benchmarks
LibriCount is a synthetic dataset for speaker count estimation.
2 papers · 0 benchmarks
Despite impressive advancements in video understanding, most efforts remain limited to coarse-grained or visual-only video tasks.
2 papers · 0 benchmarks
LuViRA (Lund University Vision, Radio, and Audio)
The Lund University Vision, Radio, and Audio (LuViRA) positioning dataset consists of 89 trajectories that are recorded in the Lund University Humanities Lab's Motion Capture (Mocap) Studio using a MIR200 robot as the targeted platform.
2 papers · 0 benchmarks
Lyra Dataset (A Dataset for Greek Traditional and Folk Music)
Lyra is a dataset of 1570 traditional and folk Greek music pieces that includes audio and video (timestamps and links to YouTube videos), along with annotations that describe aspects of particular interest for this dataset, including…
2 papers · 0 benchmarks
MVSep is a synthetic dataset for the vocal separation task created by combining random vocal and instrumental samples, publicly available on the internet.
2 papers · 0 benchmarks
MedleyVox is an evaluation dataset for multiple singing voices separation that corresponds to such categories.
2 papers · 0 benchmarks
MultiOOD (Multimodal Out-of-Distribution Detection Benchmark)
MultiOOD is the first benchmark for Multimodal OOD Detection and covers diverse dataset sizes and modalities.
2 papers · 0 benchmarks
The MuseScore dataset is a collection of 344,166 audio and MIDI pairs downloaded from MuseScore website.
2 papers · 0 benchmarks
Music4All-Onion is a large-scale, multi-modal music dataset that expands the Music4All dataset by including 26 additional audio, video, and metadata features for 109,269 music pieces and provides a set of 252,984,396 listening records of…
2 papers · 0 benchmarks
NeuroVoz (NeuroVoz: a Castillian Spanish corpus of parkinsonian speech)
The NeuroVoz dataset emerges as a pioneering resource in the field of computational linguistics and biomedical research, specifically designed to enhance the diagnosis and understanding of Parkinson's Disease (PD) through speech analysis.
2 papers · 0 benchmarks
A dataset containing the results of a MUSHRA listening test conducted with expert listeners from 2 international laboratories.
2 papers · 1 benchmark
OpenBMAT (Open Broadcast Media Audio from TV)
Open Broadcast Media Audio from TV (OpenBMAT) is an open, annotated dataset for the task of music detection that contains over 27 hours of TV broadcast audio from 4 countries distributed over 1647 one-minute long excerpts.
2 papers · 0 benchmarks
OpenSLR (Open Speech and Language Resources)
OpenSLR is a repository of open speech and language resources, including large-scale transcribed audio corpora and related software.
2 papers · 1 benchmark
Introduction The 2016 PhysioNet/CinC Challenge aims to encourage the development of algorithms to classify heart sound recordings collected from a variety of clinical or nonclinical (such as in-home visits) environments.
2 papers · 0 benchmarks
RIR dataset (Planar Room Impulse Response Dataset - ACT, DTU Electro (b. 355 r. 008))
Dataset of Room Impulse Responses measured at the Acoustic Technology group facilities, DTU Electro.
2 papers · 0 benchmarks
This repository contains the SINGA:PURA dataset, a strongly-labelled polyphonic urban sound dataset with spatiotemporal context.
2 papers · 0 benchmarks
Spatial LibriSpeech is spatial audio dataset with over 650 hours of 19-channel audio, first-order ambisonics, and optional distractor noise.
2 papers · 0 benchmarks
The speech accent archive uniformly presents a large set of speech samples from a variety of language backgrounds.
2 papers · 1 benchmark
Stanford-ECM is an egocentric multimodal dataset which comprises about 27 hours of egocentric video augmented with heart rate and acceleration data.
2 papers · 0 benchmarks
The SynthSOD dataset contains more than 47 hours of multitrack music obtained by synthesizing orchestra and ensemble pieces from the Symbolic Orchestral Database (SOD) using Spitfire BBC Symphony Orchestra Professional Library.
2 papers · 0 benchmarks
The TAU-NIGENS Spatial Sound Events 2020 dataset contains multiple spatial sound-scene recordings, consisting of sound events of distinct categories integrated into a variety of acoustical spaces, and from multiple source directions and…
2 papers · 0 benchmarks
URBAN-SED is a dataset of 10,000 soundscapes with sound event annotations generated using the scraper library.
2 papers · 0 benchmarks
VOICe is a dataset for the development and evaluation of domain adaptation methods for sound event detection.
2 papers · 0 benchmarks
VTC (Videos, Titles and Comments)
VTC is a large-scale multimodal dataset containing video-caption pairs (~300k) alongside comments that can be used for multimodal representation learning.
2 papers · 0 benchmarks
One of the founding fathers of marine mammal bioacoustics, William Watkins, carried out pioneering work with William Schevill at the Woods Hole Oceanographic Institution for more than four decades, laying the groundwork for our field today.
2 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.