Home › Datasets › modality › Speech

Speech datasets

archive 2025-07-28

197 datasets carry the modality tag "Speech", ordered by the archive's paper count. Page 1 of 5: 48 shown of 197. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Speech datasets 1–48 of 197

The LibriSpeech corpus is a collection of approximately 1,000 hours of audiobooks that are a part of the LibriVox project.
2,361 papers · 4 benchmarks
Speech Commands is an audio dataset of spoken words designed to help train and evaluate keyword spotting systems .
392 papers · 4 benchmarks
LibriTTS is a multi-speaker English corpus of approximately 585 hours of read English speech at 24kHz sampling rate, prepared by Heiga Zen with the assistance of Google Speech and Google Brain team members.
257 papers · 1 benchmark
MUSAN is a corpus of music, speech and noise.
204 papers · 0 benchmarks
AISHELL-1 is a corpus for speech recognition research and building speech recognition systems for Mandarin.
197 papers · 1 benchmark
WSJ0-2mix is a speech recognition corpus of speech mixtures using utterances from the Wall Street Journal (WSJ0) corpus.
159 papers · 3 benchmarks
LibriMix is an open-source alternative to wsj0-2mix.
122 papers · 1 benchmark
WHAM! (WSJ0 Hipster Ambient Mixtures)
The WSJ0 Hipster Ambient Mixtures (WHAM!) dataset pairs each two-speaker mixture in the wsj0-2mix dataset with a unique noise background scene.
114 papers · 2 benchmarks
VoxPopuli is a large-scale multilingual corpus providing 100K hours of unlabelled speech data in 23 languages.
113 papers · 1 benchmark
GigaSpeech, an evolving, multi-domain English speech recognition corpus with 10,000 hours of high quality labeled audio suitable for supervised training, and 40,000 hours of total audio suitable for semi-supervised and unsupervised…
87 papers · 3 benchmarks
AISHELL-2 contains 1000 hours of clean read-speech data from iOS is free for academic usage.
75 papers · 4 benchmarks
Continuous speech separation (CSS) is an approach to handling overlapped speech in conversational audio signals.
73 papers · 2 benchmarks
Multimodal Opinionlevel Sentiment Intensity (MOSI) contains: (1) multimodal observations including transcribed speech and visual gestures as well as automatic audio and visual features, (2) opinion-level subjectivity segmentation, (3)…
69 papers · 1 benchmark
CN-Celeb is a large-scale speaker recognition dataset collected in the wild'.
68 papers · 1 benchmark
CoVoST2 (Common Voice Speech-To-Text 2)
End-to-end speech-to-text translation (ST) has recently witnessed an increased interest given its system simplicity, lower inference latency and less compounding errors compared to cascaded ST (i.e.
66 papers · 0 benchmarks
ESD (Emotional Speech Database)
ESD is an Emotional Speech Database for voice conversion research.
63 papers · 0 benchmarks
VOCASET is a 4D face dataset with about 29 minutes of 4D scans captured at 60 fps and synchronized audio.
60 papers · 1 benchmark
WenetSpeech is a multi-domain Mandarin corpus consisting of 10,000+ hours high-quality labeled speech, 2,400+ hours weakly labelled speech, and about 10,000 hours unlabeled speech, with 22,400+ hours in total.
58 papers · 1 benchmark
Europarl-ST is a multilingual Spoken Language Translation corpus containing paired audio-text samples for SLT from and into 9 European languages, for a total of 72 different translation directions.
57 papers · 0 benchmarks
WHAMR! (WHAM! with synthetic reverberated sources)
WHAMR!
57 papers · 3 benchmarks
BEAT (Body-Expression-Audio-Text)
BEAT has i) 76 hours, high-quality, multi-modal data captured from 30 speakers talking with eight different emotions and in four different languages, ii) 32 millions frame-level emotion and semantic relevance annotations.
56 papers · 1 benchmark
ReVerb Challenge (REverberant Voice Enhancement and Recognition Benchmark)
The REVERB (REverberant Voice Enhancement and Recognition Benchmark) challenge is a benchmark for evaluation of automatic speech recognition techniques.
55 papers · 1 benchmark
VOICES (Voices Obscured In Complex Environmental Settings)
The VOICES corpus is a dataset to promote speech and signal processing research of speech recorded by far-field microphones in noisy room conditions.
48 papers · 0 benchmarks
AVSpeech is a large-scale audio-visual dataset comprising speech clips with no interfering background signals.
42 papers · 0 benchmarks
CHiME-5 (CHiME Speech Separation and Recognition Challenge)
The CHiME challenge series aims to advance robust automatic speech recognition (ASR) technology by promoting research at the interface of speech and language processing, signal processing , and machine learning.
42 papers · 0 benchmarks
Samanantar is the largest publicly available parallel corpora collection for Indic languages: Assamese, Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Oriya, Punjabi, Tamil, Telugu.
40 papers · 0 benchmarks
AISHELL-3 is a large-scale and high-fidelity multi-speaker Mandarin speech corpus which could be used to train multi-speaker Text-to-Speech (TTS) systems.
38 papers · 0 benchmarks
CoVoST is a large-scale multilingual speech-to-text translation corpus.
34 papers · 0 benchmarks
The DIHARD II development and evaluation sets draw from a diverse set of sources exhibiting wide variation in recording equipment, recording environment, ambient noise, number of speakers, and speaker demographics.
34 papers · 1 benchmark
THCHS-30 is a free Chinese speech database THCHS-30 that can be used to build a full-fledged Chinese speech recognition system.
34 papers · 0 benchmarks
2000 HUB5 English Evaluation Transcripts was developed by the Linguistic Data Consortium (LDC) and consists of transcripts of 40 English telephone conversations used in the 2000 HUB5 evaluation sponsored by NIST (National Institute of…
33 papers · 2 benchmarks
TIMIT (TIMIT Acoustic-Phonetic Continuous Speech Corpus)
The TIMIT Acoustic-Phonetic Continuous Speech Corpus is a standard dataset used for evaluation of automatic speech recognition systems.
31 papers · 6 benchmarks
RAVDESS (Ryerson Audio-Visual Database of Emotional Speech and Song)
The Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS) contains 7,356 files (total size: 24.8 GB).
27 papers · 6 benchmarks
CVSS is a massively multilingual-to-English speech to speech translation (S2ST) corpus, covering sentence-level parallel S2ST pairs from 21 languages into English.
26 papers · 1 benchmark
In SpokenSQuAD, the document is in spoken form, the input question is in the form of text and the answer to each question is always a span in the document.
24 papers · 1 benchmark
The Easy Communications (EasyCom) dataset is a world-first dataset designed to help mitigate the cocktail party effect from an augmented-reality (AR) -motivated multi-sensor egocentric world view.
22 papers · 4 benchmarks
SLUE (Spoken Language Understanding Evaluation)
Spoken Language Understanding Evaluation (SLUE) is a suite of benchmark tasks for spoken language understanding evaluation.
22 papers · 3 benchmarks
SEP-28k (Stuttering Events in Podcasts)
Stuttering Events in Podcasts (SEP-28k) is a dataset containing over 28k clips labeled with five event types including blocks, prolongations, sound repetitions, word repetitions, and interjections.
21 papers · 0 benchmarks
PartialSpoof is a dataset of partially-spoofed data to evaluate detection of partially-spoofed speech data.
17 papers · 0 benchmarks
The Switchboard-1 Telephone Speech Corpus (LDC97S62) consists of approximately 260 hours of speech and was originally collected by Texas Instruments in 1990-1, under DARPA sponsorship.
17 papers · 1 benchmark
SONAR, a new multilingual and multimodal fixed-size sentence embedding space, with a full suite of speech and text encoders and decoders.
16 papers · 0 benchmarks
SPGISpeech (pronounced “speegie-speech”) is a large-scale transcription dataset, freely available for academic research.
16 papers · 1 benchmark
MaSS (Multilingual corpus of Sentence-aligned Spoken utterances) is an extension of the CMU Wilderness Multilingual Speech Dataset, a speech dataset based on recorded readings of the New Testament.
15 papers · 3 benchmarks
PromptSpeech is a dataset that consists of speech and the corresponding prompts.
14 papers · 0 benchmarks
speechocean762 is an open-source speech corpus designed for pronunciation assessment use, consisting of 5000 English utterances from 250 non-native speakers, where half of the speakers are children.
14 papers · 3 benchmarks
GUM (Georgetown University Multilayer corpus)
GUM is an open source multilayer English corpus of richly annotated texts from twelve text types.
13 papers · 1 benchmark
JVS is a Japanese multi-speaker voice corpus which contains voice data of 100 speakers in three styles (normal, whisper, and falsetto).
13 papers · 0 benchmarks
Earnings-21, a 39-hour corpus of earnings calls containing entity-dense speech from nine different financial sectors.
12 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.