Home › Datasets › task › Speech Enhancement
Speech Enhancement datasets
archive 2025-07-28
24 datasets carry the task tag "Speech Enhancement" (the task itself: Speech Enhancement), ordered by the archive's paper count. Page 1 of 1: 24 shown of 24. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Speech Enhancement datasets 1–24 of 24
LibriMix is an open-source alternative to wsj0-2mix.
122 papers · 1 benchmark
WHAM! (WSJ0 Hipster Ambient Mixtures)
The WSJ0 Hipster Ambient Mixtures (WHAM!) dataset pairs each two-speaker mixture in the wsj0-2mix dataset with a unique noise background scene.
114 papers · 2 benchmarks
AVA (Atomic Visual Actions)
AVA is a project that provides audiovisual annotations of video for improving our understanding of human activity.
113 papers · 7 benchmarks
WHAMR! (WHAM! with synthetic reverberated sources)
WHAMR!
57 papers · 3 benchmarks
The REVERB (REverberant Voice Enhancement and Recognition Benchmark) challenge is a benchmark for evaluation of automatic speech recognition techniques.
55 papers · 1 benchmark
VoiceBank + DEMAND (Noisy speech database for training speech enhancement algorithms and TTS models)
VoiceBank+DEMAND is a noisy speech database for training speech enhancement algorithms and TTS models.
53 papers · 1 benchmark
The DNS Challenge at INTERSPEECH 2020 intended to promote collaborative research in single-channel Speech Enhancement aimed to maximize the perceptual quality and intelligibility of the enhanced speech.
48 papers · 3 benchmarks
CHiME-5 (CHiME Speech Separation and Recognition Challenge)
The CHiME challenge series aims to advance robust automatic speech recognition (ASR) technology by promoting research at the interface of speech and language processing, signal processing , and machine learning.
42 papers · 0 benchmarks
TIMIT (TIMIT Acoustic-Phonetic Continuous Speech Corpus)
The TIMIT Acoustic-Phonetic Continuous Speech Corpus is a standard dataset used for evaluation of automatic speech recognition systems.
31 papers · 6 benchmarks
Contains temporally labeled face tracks in video, where each face instance is labeled as speaking or not, and whether the speech is audible.
22 papers · 1 benchmark
The Easy Communications (EasyCom) dataset is a world-first dataset designed to help mitigate the cocktail party effect from an augmented-reality (AR) -motivated multi-sensor egocentric world view.
22 papers · 4 benchmarks
VoiceBank+DEMAND is a noisy speech database for training speech enhancement algorithms and TTS models.
16 papers · 1 benchmark
L3DAS22: MACHINE LEARNING FOR 3D AUDIO SIGNAL PROCESSING This dataset supports the L3DAS22 IEEE ICASSP Gand Challenge.
13 papers · 0 benchmarks
The EARS-WHAM dataset mixes speech from the EARS dataset with real noise recordings from the WHAM!
11 papers · 1 benchmark
The QMUL underGround Re-IDentification (GRID) dataset contains 250 pedestrian image pairs.
10 papers · 5 benchmarks
L3DAS21 is a dataset for 3D audio signal processing.
6 papers · 2 benchmarks
RealMAN (A Real-Recorded and Annotated Microphone Array Dataset for Dynamic Speech Enhancement and Localization)
The Audio Signal and Information Processing Lab at Westlake University, in collaboration with AISHELL, has released the Real-recorded and annotated Microphone Array speech&Noise (RealMAN) dataset, which provides annotated multi-channel…
5 papers · 2 benchmarks
A new large-scale, in-thewild Mandarin dataset, CAS-VSR-S101 with 101.1 hours of data.
1 paper · 3 benchmarks
A Brazilian Portuguese TTS dataset featuring a female voice recorded with high quality in a controlled environment, with neutral emotion and more than 20 hours of recordings.
1 paper · 0 benchmarks
A database containing high sampling rate recordings of a single speaker reading sentences in Brazilian Portuguese with neutral voice, along with the corresponding text corpus.
1 paper · 0 benchmarks
The NISQA Corpus includes more than 14,000 speech samples with simulated (e.g.
1 paper · 0 benchmarks
Uses same clean speech as VoiceBank+Demand but more noise types.
1 paper · 1 benchmark
WHAMRext is an extension to the WHAMR corpus with larger RT60 values (between 1s and 3s)
1 paper · 1 benchmark
mDRT (Multilingual Diagnostic Rhyme Test)
We present a multilingual test set for conducting speech intelligibility tests in the form of diagnostic rhyme tests.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.