Home › Datasets › modality › Audio

Audio datasets

archive 2025-07-28

480 datasets carry the modality tag "Audio", ordered by the archive's paper count. Page 10 of 10: 48 shown of 480. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Audio datasets 433–480 of 480

This open-source dataset consists of 5.04 hours of transcribed English conversational speech beyond telephony, where 13 conversations were contained.
0 papers · 0 benchmarks
The Arabic Speech Corpus (1.5 GB) is a Modern Standard Arabic (MSA) speech corpus for speech synthesis.
0 papers · 0 benchmarks
AudioSet CC (AudioSet Creative Commons)
The subset of audio samples from the AudioSet ontology which are licensed with Creative Commons.
0 papers · 0 benchmarks
BAVL (Blind Audio-Visual Localization (BAVL))
Blind Audio-Visual Localization (BAVL) Dataset consists of 20 audio-visual recordings of sound sources, which could be talking faces or music instruments.
0 papers · 0 benchmarks
Bach chorales is a univariate time series based on chorales, where the task is to learn generative grammar.
0 papers · 0 benchmarks
The BirdVox-DCASE-20k dataset contains 20,000 ten-second audio recordings.
0 papers · 0 benchmarks
CLO-43SD is a dataset for multi-class species identification in avian flight calls.
0 papers · 0 benchmarks
CLO-SWTH is a dataset for species-specific flight call identification for the Swainson’s Thrush.
0 papers · 0 benchmarks
CLO-WTSP is a dataset for species-specific flight call identification for the White-Throated Sparrow.
0 papers · 0 benchmarks
This publicly available data is synthesised audio for woodwind quartets including renderings of each instrument in isolation.
0 papers · 0 benchmarks
Chernobyl is a collection of 620 audio clips collected from unattended remote monitoring equipment in the Chernobyl Exclusion Zone (CEZ).
0 papers · 0 benchmarks
Couples Therapy (Couples Therapy Corpus)
The Couples Therapy corpus contains audio, video recordings and manual transcriptions of conversations between 134 real-life couples attending marital therapy.
0 papers · 0 benchmarks
DBR dataset is an environmental audio dataset created for the Bachelor's Seminar in Signal Processing in Tampere University of Technology.
0 papers · 0 benchmarks
Dataset Summary The Deep Evaluation of Audio Representations (DEAR) dataset is a benchmark designed to assess general-purpose audio foundation models on properties critical for hearable devices.
0 papers · 0 benchmarks
Deeply vocal characterizer is a human nonverbal vocalization dataset.
0 papers · 0 benchmarks
DementiaBank is a shared database of multimedia interactions for the study of communication in dementia.
0 papers · 0 benchmarks
FSL4 (Freesound Loops 4k)
The FSL4 dataset contains ~4000 user-contributed loops uploaded to Freesound.
0 papers · 0 benchmarks
Fongbe Speech Dataset (Fongbe speech dataset V2)
This dataset was created for Fongbe automatic speech recognition task and contains about 3979 recordings of 13 participants reading a text written in Fongbe, one sentence at a time.
0 papers · 0 benchmarks
We introduce FortisAVQA, a dataset designed to assess the robustness of AVQA models.
0 papers · 0 benchmarks
HEADSET (HEADSET: Human Emotion Awareness under Partial Occlusions Multimodal DataSET)
The volumetric representation of human interactions is one of the fundamental domains in the development of immersive media productions and telecommunication applications.
0 papers · 0 benchmarks
ISMIR Genre (ISMIR2004 Genre)
ISMIR2004 is an audio dataset consisting of 6 genres with 729 excerpts of 30 seconds.
0 papers · 0 benchmarks
MAEC (Multimodal Aligned Earnings Conference Call Dataset)
MAEC is a new, large-scale multi-modal, text-audio paired, earnings-call dataset named MAEC, based on S&P 1500 companies.
0 papers · 0 benchmarks
MCCSD (Mandarin Chinese Cued Speech Dataset)
This MCCS dataset is the first large-scale Mandarin Chinese Cued Speech dataset.
0 papers · 0 benchmarks
MHRI dataset (Multimodal Human-Robot Interaction dataset)
The dataset includes recordings from 10 different users teaching the robot different common kitchen objects, that consists of synchronized recordings from three cameras and a microphone mounted on the robot: An RGB-d camera covers the user…
0 papers · 0 benchmarks
The MOBIO database consists of bi-modal (audio and video) data taken from 152 people.
0 papers · 0 benchmarks
MedleyDB 2.0 is a superset of the MedleyDB – a dataset of annotated, royalty-free multitrack recordings.
0 papers · 0 benchmarks
The MIVIA audio events data set is composed of a total of 6000 events for surveillance applications, namely glass breaking, gun shots and screams.
0 papers · 0 benchmarks
Mixing Secrets is an instrument recognition dataset containing 258 multi-track recordings sourced from the Mixing Secrets for The Small Studio website.
0 papers · 0 benchmarks
Mudestreda (Mudestreda Multimodal Device State Recognition Dataset)
Mudestreda Multimodal Device State Recognition Dataset obtained from real industrial milling device with Time Series and Image Data for Classification, Regression, Anomaly Detection, Remaining Useful Life (RUL) estimation, Signal Drift…
0 papers · 0 benchmarks
PCVC (Persian Consonant Vowel Combination)
The Persian Consonant Vowel Combination (PCVC) dataset is a phoneme based speech dataset, and also the first free Persian speech dataset to help Persian speech researchers.
0 papers · 0 benchmarks
PolandNFC is a collection of 4,000 recordings from Hanna Pamuła's PhD project of monitoring autumn nocturnal bird migration.
0 papers · 0 benchmarks
Robbie Williams is a dataset of 65 songs by Robbie Williams.
0 papers · 0 benchmarks
STVD-FC (Fact-checking dataset)
STVD-FC is the largest public dataset on the political content analysis and fact-checking tasks.
0 papers · 0 benchmarks
Saarbruecken Voice Database contains voice and EGG recordings of patients diagnosed with voice disorder, as well as healthy persons.
0 papers · 0 benchmarks
SimSceneTVB is a dataset of 600 simulated sound scenes of 45s each representing urban sound environments, simulated using the simScene Matlab library.
0 papers · 0 benchmarks
SimSceneTVB Perception is a corpus of 100 sound scenes of 45s each representing urban sound environments, including: 6 scenes recorded in Paris, 19 scenes simulated using simScene to replicate recorded scenarios, 75 scenes simulated using…
0 papers · 0 benchmarks
The Sound Events for Surveillance Applications (SESA) dataset files were obtained from Freesound.
0 papers · 0 benchmarks
The TAU Spatial Sound Events 2019 - Ambisonic dataset contains recordings from a scene (along with the Microphone Array sister dataset).
0 papers · 0 benchmarks
The TAU Spatial Sound Events 2019 – Microphone Array dataset contains recordings from a scene (along with the Ambisonic sister dataset).
0 papers · 0 benchmarks
THVD (Talking Head Video Dataset)
About We provide a comprehensive talking-head video dataset with over 50,000 videos, totaling more than 500+ hours of footage and featuring 20,841 unique identities from around the world.
0 papers · 0 benchmarks
TUT Rare Sound events 2017, development dataset consists of source files for creating mixtures of rare sound events (classes baby cry, gun shot, glass break) with background audio, as well a set of readily generated mixtures and recipes…
0 papers · 0 benchmarks
The TUT Sounds Event 2018 dataset consists of real-life first order Ambisonic (FOA) format recordings with stationary point sources each associated with a spatial coordinate.
0 papers · 0 benchmarks
Video Dataset (Storytelling Video Dataset (Russian, Emotion, Gesture, Speech))
The Storytelling Video Dataset is a high-quality, human-reviewed multimodal dataset featuring over 700 full-body video recordings of native Russian speakers.
0 papers · 0 benchmarks
VocSim (Vocal Similarity Benchmark)
VocSim (Vocal Similarity Benchmark) is a benchmark designed to evaluate the ability of neural audio embeddings to capture acoustic and perceptual similarity in a zero-shot setting, without task-specific fine-tuning.
0 papers · 0 benchmarks
The WASABI Song Corpus is a large corpus of songs enriched with metadata extracted from music databases on the Web, and resulting from the processing of song lyrics and from audio analysis.
0 papers · 0 benchmarks
Yesno is an audio dataset consisting of 60 recordings of one individual saying yes or no in Hebrew; each recording is eight words long.
0 papers · 0 benchmarks
Freefield1010 is a collection of 7,690 excerpts from field recordings around the world, gathered by the FreeSound project, and then standardised for research.
0 papers · 0 benchmarks
warblrb10k is a collection of 10,000 smartphone audio recordings from around the UK, crowdsourced by users of Warblr the bird recognition app.
0 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.