Home › Datasets › modality › Audio
Audio datasets
archive 2025-07-28
480 datasets carry the modality tag "Audio", ordered by the archive's paper count. Page 7 of 10: 48 shown of 480. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Audio datasets 289–336 of 480
WildDESED (Wild Domestic Environment Sound Event Detection)
WildDESED is an extension of the original DESED dataset, created to reflect various domestic scenarios by incorporating complex and unpredictable background noises.
2 papers · 1 benchmark
We present YTSeg, a topically and structurally diverse benchmark for the text segmentation task based on YouTube transcriptions.
2 papers · 1 benchmark
The YouTube8M-MusicTextClips dataset consists of over 4k high-quality human text descriptions of music found in video clips from the YouTube8M dataset.
2 papers · 0 benchmarks
Overview nEMO is a simulated dataset of emotional speech in the Polish language.
2 papers · 0 benchmarks
3D-Speaker is a large-scale speech corpus designed to facilitate the research of speech representation disentanglement.
1 paper · 0 benchmarks
AESI (Athens Emotional States Inventory)
The development of ecologically valid procedures for collecting reliable and unbiased emotional data towards computer interfaces with social and affective intelligence targeting patients with mental disorders.
1 paper · 0 benchmarks
ARTE (Ambisonics Recordings of Typical Environments)
The ARTE database, so far, contains 13 acoustic environments that were recorded with a purpose-built 62-channel microphone array in various locations around Sydney (Australia), and was decoded into the higher-order Ambisonics (HOA) format.
1 paper · 0 benchmarks
ARVSU (Addressee Recognition in Visual Scenes with Utterances)
ARVSU contains a vast body of image variations in visual scenes with an annotated utterance and a corresponding addressee for each scenario.
1 paper · 0 benchmarks
A Rich Annotated Mandarin Conversational (RAMC) Speech Dataset, including 180 hours of Mandarin Chinese dialogue, 150, 10 and 20 hours for the training set, development set and test set respectively.
1 paper · 0 benchmarks
ATD-Dataset (Auto-Tune Detection Dataset (ATD-Dataset))
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
AVASpeech-SMAD (AVASpeech-SMAD: A Strongly Labelled Speech and Music Activity Detection Dataset with Label Co-Occurrence)
We propose a dataset, AVASpeech-SMAD, to assist speech and music activity detection research.
1 paper · 0 benchmarks
Provided in the linked paper.
1 paper · 0 benchmarks
AViMoS (Audio-Visual Mouse Saliency)
A novel audio-visual mouse saliency (AViMoS) dataset with the following key-features: Diverse content: movie, sports, live, vertical videos, etc.; Large scale: 1500 videos with mean 19s duration; High resolution: all streams are FullHD;…
1 paper · 0 benchmarks
AcousticRooms is a large-scale synthetic room impulse response (RIR) dataset designed for cross-room RIR prediction tasks.
1 paper · 0 benchmarks
The AllMusic Mood Subset (AMS) is a dataset for mood classification from songs.
1 paper · 0 benchmarks
ArVoice (ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis)
We introduce ArVoice, a multi-speaker Modern Standard Arabic (MSA) speech corpus with diacritized transcriptions, intended for multi-speaker speech synthesis, and can be useful for other tasks such as speech-based diacritic restoration,…
1 paper · 0 benchmarks
Dataset Description: The dataset comprises audio recordings of the wing beats of Aedes aegypti mosquitoes and others, conducted in a semi-controlled environment.
1 paper · 0 benchmarks
Test dataset for unsupervised anomaly detection in sound (ADS).
1 paper · 0 benchmarks
Temporal Dataset for Indoor and In-Vehicle Thermal Comfort Estimation Abstract Thermal comfort estimation is essential for enhancing user experience in static indoor environments and dynamic in-vehicle scenarios.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
BAH (Behavioural Ambivalence/Hesitancy)
Recognizing complex emotions linked to ambivalence and hesitancy (A/H) can play a critical role in the personalization and effectiveness of digital behaviour change interventions.
1 paper · 0 benchmarks
BERSt (Basic Emotion Random phrase Shouts)
BERSt Dataset We release the BERSt Dataset for various speech recognition tasks including Automatic Speech Recognition (ASR) and Speech Emotion Recogniton (SER) Overview 4526 single phrase recordings (~3.75h) 98 professional actors 19…
1 paper · 1 benchmark
We recorded gun sounds by changing the type and position of guns to diversify distances and angles in the PUBG environment.
1 paper · 0 benchmarks
BIOSED-ACPD: BIOacoustic Sound Event Detection - Adaptive Change Point Detection dataset Description.
1 paper · 0 benchmarks
BIRD (Big Impulse Response Dataset) is an open dataset that consists of 100,000 multichannel room impulse responses (RIRs) generated from simulations using the Image Method, making it the largest multichannel open dataset currently…
1 paper · 0 benchmarks
The BIRDeep Audio Annotations dataset is a collection of bird vocalizations from Doñana National Park, Spain.
1 paper · 0 benchmarks
A Bilingual Dataset for Bangla and English Voice Commands Colloquial Bangla has adopted many English words due to colonial influence.
1 paper · 1 benchmark
In this dataset two robots, Baxter and UR5, perform 8 behaviors (look, grasp, pick, hold, shake, lower, drop, and push) on 95 objects that vary by 5 color (blue, green, red, white, and yellow), 6 contents (wooden button, plastic dices,…
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Boombox is a multi-modal dataset for visual reconstruction from acoustic vibrations.
1 paper · 0 benchmarks
CAL10K (Computer Audition Lab 10000)
The CAL10K dataset (introduced as Swat10k) contains 10,870 songs that are weakly-labelled using a tag vocabulary of 475 acoustic tags and 153 genre tags.
1 paper · 0 benchmarks
The CAL500 Expansion (CAL500exp) dataset is an enriched version of the CAL500 music information retrieval dataset.
1 paper · 0 benchmarks
A new large-scale, in-thewild Mandarin dataset, CAS-VSR-S101 with 101.1 hours of data.
1 paper · 3 benchmarks
CHORD (CHOrus Recognition Dataset)
CHORD is the first chorus recognition dataset containing 627 songs for public use.
1 paper · 0 benchmarks
CUCO Database (A voice and speech corpus of patients who underwent upper airway surgery in pre-and post-operative states)
Many research articles have explored the impact of surgical interventions on voice and speech evaluations, but advances are limited by the lack of publicly accessible datasets.
1 paper · 0 benchmarks
In this dataset an uppertorso humanoid robot with 7-DOF arm explored 100 different objects belonging to 20 different categories using 10 behaviors: Look, Crush, Grasp, Hold, Lift, Drop, Poke, Push, Shake and Tap.
1 paper · 0 benchmarks
35 recordings of Candombe music with beat and downbeat annotations.
1 paper · 2 benchmarks
ChMusic is a traditional Chinese music dataset for training model and performance evaluation of musical instrument recognition.
1 paper · 0 benchmarks
The Cleft dataset is a collection of ultrasound tongue imaging and audio data, gathered from children with cleft lip and palate by a research speech and language therapist working in a hospital environment.
1 paper · 0 benchmarks
We construct a large-scale conducting motion dataset, named ConductorMotion100, by deploying pose estimation on conductor view videos of concert performance recordings collected from online video platforms.
1 paper · 0 benchmarks
- Mood ratings of 8 emotions gathered across 360 pop songs - 166 raters from US, S.Korea and Brazil - MIR features from Spotify
1 paper · 0 benchmarks
DINOS (Diverse INdustrial Operation Sounds)
DINOS (Diverse INdustrial Operation Sounds) is a large-scale, open-access dataset consisting of over 74,000 audio samples totaling more than 1,093 hours, collected from a wide range of industrial acoustic scenarios.
1 paper · 0 benchmarks
This is a dataset of 22.5 hours of synthesized audio using the open-source learnfm clone of the DX7 FM synthesizer, based upon 31K presets from Bobby Blue.
1 paper · 0 benchmarks
We release both the processed data and evaluation results from our own experiments, and the underlying raw data that can be used for future experiments and schemes in the domain of Zero-Interaction Security.
1 paper · 0 benchmarks
This dataset contains recordings of 32 sound producing insect species with a total 335 files and a length of 57 minutes.
1 paper · 0 benchmarks
DnR-nonverbal is a dataset for cinematic audio source separation (CASS) based on Divide and Remaster (DnR) dataset.
1 paper · 0 benchmarks
This dataset contains transcriptions of the electric guitar performance of 240 tablatures, rendered with different tones.
1 paper · 0 benchmarks
The ENF moving video dataset, which is a subset of the dataset used in Temporal Localization of Non-Static Digital Videos Using the Electrical Network Frequency , consists of video recording without the audio channel coupled with the…
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.