Home › Datasets › modality › Audio
Audio datasets
archive 2025-07-28
480 datasets carry the modality tag "Audio", ordered by the archive's paper count. Page 8 of 10: 48 shown of 480. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Audio datasets 337–384 of 480
EmoFilm (Emotional speech from Films)
EmoFilm is a multilingual emotional speech corpus comprising 1115 audio instances produced in English, Italian, and Spanish languages.
1 paper · 0 benchmarks
EmoSpeech contains keywords with diverse emotions and background sounds, presented to explore new challenges in audio analysis.
1 paper · 0 benchmarks
ErhuPT (Erhu Playing Technique Dataset)
This dataset is an audio dataset containing about 1500 audio clips recorded by multiple professional players.
1 paper · 0 benchmarks
FSC-P2 (Fearless Steps Challenge Phase2)
The Fearless Steps Initiative by UTDallas-CRSS led to the digitization, recovery, and diarization of 19,000 hours of original analog audio data, as well as the development of algorithms to extract meaningful information from this…
1 paper · 0 benchmarks
Fongbe Data collected by Fréjus A.
1 paper · 1 benchmark
The Freesound One-Shot Percussive Sounds dataset contains 10254 one-shot (single event) percussive sounds from Freesound.org and the corresponding timbral analysis.
1 paper · 0 benchmarks
Freiburg Terrains consist of three parts: 3.7 hours of audio recordings of the microphone pointed at the robot wheels.
1 paper · 0 benchmarks
GPLA-12 is a new acoustic leakage dataset of gas pipelines involving 12 categories over 684 training/testing acoustic signals.
1 paper · 0 benchmarks
A Brazilian Portuguese TTS dataset featuring a female voice recorded with high quality in a controlled environment, with neutral emotion and more than 20 hours of recordings.
1 paper · 0 benchmarks
A database containing high sampling rate recordings of a single speaker reading sentences in Brazilian Portuguese with neutral voice, along with the corresponding text corpus.
1 paper · 0 benchmarks
Guitar-TECHS (Guitar Tones/Techniques, Excerpts & Chords Dataset)
Guitar-TECHS is a comprehensive dataset featuring a variety of guitar techniques, musical excerpts, chords, and scales.
1 paper · 0 benchmarks
Beats, downbeats, and functional structural annotations for 912 Pop tracks.
1 paper · 2 benchmarks
The Haydn Annotation Dataset consists of note onset annotations from 24 experiment participants with varying musical experience.
1 paper · 0 benchmarks
IMaSC (ICFOSS Malayalam Speech Corpus)
IMaSC is a Malayalam text and speech corpus made available by ICFOSS for the purpose of developing speech technology for Malayalam, particularly text-to-speech.
1 paper · 0 benchmarks
We introduce a new synthetic test set named IS3 for interactive sound source localization.
1 paper · 0 benchmarks
JAAH (Jazz Audio-Aligned Harmony)
Eremenko, E.
1 paper · 2 benchmarks
📊 Dataset Details - Name: JamendoMaxCaps - URL: https://huggingface.co/datasets/amaai-lab/JamendoMaxCaps - Content: 362,238 songs with captions generated by Qwen2-Audio Metadata Fields - genre - speed - variable tags 🎯 Rationale 1.
1 paper · 0 benchmarks
Kinect-WSJ is a multichannel, multispeaker, reverberated, noisy dataset which extends the WSJ0-2mix singlechannel, non-reverberated, noiseless dataset to the strong reverberation and noise conditions and the Kinect-like microphone array…
1 paper · 0 benchmarks
Introduction These audio files accompany the preprint by Accolti (2025), which presents a preliminary study on the effect of the acoustical conditions of three different rooms on the perception of virtual stages for music.
1 paper · 0 benchmarks
This dataset contains two types of audio recordings.
1 paper · 0 benchmarks
The M-AILABS Speech Dataset is the first large dataset that we are providing free-of-charge, freely usable as training data for speech recognition and speech synthesis.
1 paper · 1 benchmark
Characterising multimedia content with relevant, reliable and discriminating tags is vital for multimedia information retrieval.
1 paper · 0 benchmarks
Here we release the dataset (MultiChannelGrid, abbreviated as MCGrid) used in our paper LIMUSE: LIGHTWEIGHT MULTI-MODAL SPEAKER EXTRACTION](https://arxiv.org/abs/2111.04063)).
1 paper · 0 benchmarks
MINT (a Multi-modal Image and Narrative Text Dubbing Dataset)
Foley audio, critical for enhancing the immersive experience in multimedia content, faces significant challenges in the AI-generated content (AIGC) landscape.
1 paper · 0 benchmarks
Periodic Tic sounds (T0=1s) sampled at 16kHz with duration of nearly 10s.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
MediBeng (Synthetic Code-Switched Bengali-English Speech Conversations for Healthcare Applications)
MediBeng Dataset The MediBeng dataset contains synthetic code-switched dialogues in Bengali and English for training models in speech recognition (ASR), text-to-speech (TTS), and machine translation in clinical settings.
1 paper · 1 benchmark
A dataset called Medley2K that consists of 2,000 medleys and 7,712 labeled transitions.
1 paper · 0 benchmarks
A large-scale reference dataset for bioacoustics.
1 paper · 1 benchmark
Dataset for multimodal skills assessment focusing on assessing piano player’s skill level.
1 paper · 3 benchmarks
MuseChat Dataset (MuseChat: A Conversational Music Recommendation System for Videos (CVPR 2024 Highlight Paper))
Music recommendation for videos attracts growing interest in multi-modal research.
1 paper · 0 benchmarks
NAR is a dataset of audio recordings made with the humanoid robot Nao in real world conditions for sound recognition benchmarking.
1 paper · 0 benchmarks
The NISQA Corpus includes more than 14,000 speech samples with simulated (e.g.
1 paper · 0 benchmarks
Nlakh is a dataset for Musical Instrument Retrieval.
1 paper · 0 benchmarks
OpenSpeaks Voice: Odia is a large speech dataset in the Odia language of India that is stewarded by Subhashish Panigrahi and is hosted at the O Foundation.
1 paper · 0 benchmarks
PC-GITA is a Spanish speech corpus designed to analyze speech impairments in individuals with Parkinson's Disease (PD).
1 paper · 0 benchmarks
PIAST (PIAST: A Multimodal Piano Dataset with Audio, Symbolic and Text)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
The POTUS Corpus is a Database of Weekly Addresses for the Study of Stance in Politics and Virtual Agents.
1 paper · 0 benchmarks
Parkinson Speech Dataset is an audio dataset consisting of recordings of 20 Parkinson's Disease (PD) patients and 20 healthy subjects.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
In this Pre-Contest Workshop Video Recordings folder: Seven screen and audio recordings of seven pre-contest workshops
1 paper · 0 benchmarks
Quechua Collao corpus for automatic emotion recognition in speech.
1 paper · 1 benchmark
Asthma is a common, usually long-term respiratory disease with negative impact on society and the economy worldwide.
1 paper · 0 benchmarks
The RWCP Sound Scene Database includes non-speech sounds recorded in an anechoic room, reconstructed signals in various rooms, impulse responses for a microphone array, speech data recorded with the same array, and recordings of background…
1 paper · 1 benchmark
The full version of ReefSet used in Williams et al.
1 paper · 0 benchmarks
SES (Spanish Emotional Speech)
Currently, an essential point in speech synthesis is the addressing of the variability of human speech.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.