Home › Datasets › modality › Audio

Audio datasets

archive 2025-07-28

480 datasets carry the modality tag "Audio", ordered by the archive's paper count. Page 9 of 10: 48 shown of 480. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Audio datasets 385–432 of 480

SF20K (Short-Films 20K)
Short-Films 20K (SF20K) is the largest publicly available movie dataset.
1 paper · 0 benchmarks
SHD - Adding (Spiking Heidelberg Digits - Adding)
This dataset is based on the Spiking Heidelberg Digits (SHD) dataset.
1 paper · 1 benchmark
F.
1 paper · 1 benchmark
SINS is a database of continuous real-life audio recordings in a home environment.
1 paper · 1 benchmark
A.
1 paper · 1 benchmark
SaGA (The Bielefeld Speech and Gesture Alignment Corpus (SaGA))
The primary data of the SaGA corpus are made up of 25 dialogs of interlocutors (50), who engage in a spatial communication task combining direction-giving and sight description.
1 paper · 0 benchmarks
Speech Recognition Dataset for Oromo Language.
1 paper · 1 benchmark
ShiftySpeech: A Large-Scale Synthetic Speech Dataset with Distribution Shifts 🔥 Key Features - 3000+ hours of synthetic speech - Diverse Distribution Shifts: The dataset spans 7 key distribution shifts, including: - 📖 Reading Style - 🎙️…
1 paper · 0 benchmarks
Skit-S2I (Skit-S2I: An Indian Accented Speech to Intent dataset)
This dataset for Intent classification from human speech covers 14 coarse-grained intents from the Banking domain.
1 paper · 0 benchmarks
SoccerNet-Echoes (SoccerNet-Echoes: A Soccer Game Audio Commentary Dataset)
SoccerNet-Echoes: A Soccer Game Audio Commentary Dataset.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
We collect a dataset of 805 clean videos that show the action of pouring water in a container.
1 paper · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Synthetic Speech Attribution Dataset.
1 paper · 0 benchmarks
J.
1 paper · 2 benchmarks
The SWC is a corpus of aligned Spoken Wikipedia articles from the English, German, and Dutch Wikipedia.
1 paper · 1 benchmark
Thorsten-Voice (Thorsten-21.02-neutral) is a neutrally spoken voice dataset recorded by Thorsten Müller, audio optimized by Dominik Kreutz and licenced under CC0 to provide it for anybody without any financial or licence struggle.
1 paper · 1 benchmark
TinyChirp dataset for model training, validation and testing
1 paper · 0 benchmarks
In this dataset UR5 robot used 6 tools: metal-scissor, metal-whisk, plastic-knife, plastic-spoon, wooden-chopstick, and wooden-fork to perform 6 behaviors: look, stirring-slow, stirring-fast, stirring-twist, whisk, and poke.
1 paper · 0 benchmarks
USC (Uzbek Speech Corpus)
The Uzbek speech corpus (USC) comprises 958 different speakers with a total of 105 hours of transcribed audio recordings.
1 paper · 0 benchmarks
USM-SED is a dataset for polyphonic sound event detection in urban sound monitoring use-cases.
1 paper · 0 benchmarks
We present UniTalk, a novel dataset specifically designed for the task of active speaker detection, emphasizing challenging scenarios to enhance model generalization.
1 paper · 0 benchmarks
The United-Syn-Med dataset is a specialized medical speech dataset designed to evaluate and improve Automatic Speech Recognition (ASR) systems within the healthcare domain.
1 paper · 0 benchmarks
A main goal of the Urban Soundscapes of the World project is to create a reference database of examples of urban acoustic environments, consisting of high-quality immersive audiovisual recordings (360-degree video and spatial audio), in…
1 paper · 0 benchmarks
VAST Absorption is a dataset of spatial binaural features annotated with acoustic properties such as the 3D source position and the walls’ absorption coefficients.
1 paper · 0 benchmarks
This is the reference headset microphone variant of the VibraVox dataset.
1 paper · 2 benchmarks
This is the in-ear rigid earpiece-embedded microphone variant of the VibraVox dataset.
1 paper · 3 benchmarks
This is the in-ear comply foam-embedded microphone variant of the VibraVox dataset.
1 paper · 3 benchmarks
This is the throat microphone (laryngophone) variant of the VibraVox dataset.
1 paper · 3 benchmarks
In doctor-patient conversations, identifying medically relevant information is crucial, posing the need for conversation summarization.
1 paper · 0 benchmarks
Virtuoso Strings is a dataset for soft onsets detection for string instruments.
1 paper · 0 benchmarks
The VocalImitationSet is a collection of crowd-sourced vocal imitations of a large set of diverse sounds collected from Freesound (https://freesound.org/), which were curated based on Google's AudioSet ontology…
1 paper · 0 benchmarks
VoxForge is an open speech dataset that was set up to collect transcribed speech for use with Free and Open Source Speech Recognition Engines (on Linux, Windows and Mac).
1 paper · 1 benchmark
WHAMRext is an extension to the WHAMR corpus with larger RT60 values (between 1s and 3s)
1 paper · 1 benchmark
WHU - Audio ENF (WHU - Audio Electric Network Frequency)
The Whu dataset is an audio only dataset thought for testing ENF detection.
1 paper · 0 benchmarks
Well-being Dataset (Cambridge Well-being Dataset for Psychological Distress Analysis)
The dataset is a private dataset collected for automatic analysis of psychological distress.
1 paper · 1 benchmark
Manually labelled dataset of bird recordings from the species of interest inhabiting in the wetlands of the "Aiguamolls del Empord\{a}" natural park in Girona, Spain.
1 paper · 0 benchmarks
The Western Mediterranean Wetlands Bird Dataset is a collection of birds' vocalizations of different lengths that primarily consists of 5,795 labelled audio clips derived from 1,098 recordings, totalling 201.6 minutes or 12,096 seconds…
1 paper · 0 benchmarks
XMIDI is a comprehensive, large-scale symbolic music dataset that includes accurate emotion and genre labels, consisting of 108,023 MIDI files.
1 paper · 0 benchmarks
We redistribute a suite of datasets as part of the YourMT3 project.
1 paper · 0 benchmarks
Zooniverse (HumBug Zooniverse)
The Humbug Zooinverse dataset is a dataset of mosquito audio recordings.
1 paper · 0 benchmarks
gtzanmusicspeech is a dataset for music/speech discrimination.
1 paper · 0 benchmarks
inaGVAD (InaGVAD : a Challenging French TV and Radio Corpus annotated for Voice Activity Detection and Speaker Gender Segmentation)
InaGVAD is a Voice Activity Detection (VAD) and Speaker Gender Segmentation (SGS) dataset designed for representing the acoustic diversity of French TV and Radio programs.
1 paper · 0 benchmarks
l2d (Learning to Dance)
This dataset is composed of paired videos of people dancing 3 different music styles: Ballet, Michael Jackson and Salsa.
1 paper · 0 benchmarks
mDRT (Multilingual Diagnostic Rhyme Test)
We present a multilingual test set for conducting speech intelligibility tests in the form of diagnostic rhyme tests.
1 paper · 0 benchmarks
taste-music-dataset (Taste Music Dataset)
This dataset is a patched version of The Taste & Affect Music Database by D.
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.