Home › Datasets › task › Speech Recognition

Speech Recognition datasets

archive 2025-07-28

96 datasets carry the task tag "Speech Recognition" (the task itself: Speech Recognition), ordered by the archive's paper count. Page 1 of 2: 48 shown of 96. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Speech Recognition datasets 1–48 of 96

The MNIST database (Modified National Institute of Standards and Technology database) is a large collection of handwritten digits.
7,651 papers · 44 benchmarks
The LibriSpeech corpus is a collection of approximately 1,000 hours of audiobooks that are a part of the LibriVox project.
2,361 papers · 4 benchmarks
Common Voice is an audio dataset that consists of a unique MP3 and corresponding text file.
449 papers · 143 benchmarks
Speech Commands is an audio dataset of spoken words designed to help train and evaluate keyword spotting systems .
392 papers · 4 benchmarks
MuST-C currently represents the largest publicly available multilingual corpus (one-to-many) for speech translation.
216 papers · 2 benchmarks
AISHELL-1 is a corpus for speech recognition research and building speech recognition systems for Mandarin.
197 papers · 1 benchmark
Libri-Light is a collection of spoken English audio suitable for training speech recognition systems under limited or no supervision.
194 papers · 2 benchmarks
FLEURS (Few-shot Learning Evaluation of Universal Representations of Speech)
We introduce FLEURS, the Few-shot Learning Evaluation of Universal Representations of Speech benchmark.
141 papers · 1 benchmark
Europarl (European Parliament Proceedings Parallel Corpus)
A corpus of parallel text in 21 European languages from the proceedings of the European Parliament.
128 papers · 1 benchmark
LRS2 (Lip Reading Sentences 2)
The Oxford-BBC Lip Reading Sentences 2 (LRS2) dataset is one of the largest publicly available datasets for lip reading sentences in-the-wild.
115 papers · 10 benchmarks
VoxPopuli is a large-scale multilingual corpus providing 100K hours of unlabelled speech data in 23 languages.
113 papers · 1 benchmark
GigaSpeech, an evolving, multi-domain English speech recognition corpus with 10,000 hours of high quality labeled audio suitable for supervised training, and 40,000 hours of total audio suitable for semi-supervised and unsupervised…
87 papers · 3 benchmarks
Multilingual LibriSpeech is a large multilingual corpus suitable for speech research.
77 papers · 2 benchmarks
AISHELL-2 contains 1000 hours of clean read-speech data from iOS is free for academic usage.
75 papers · 4 benchmarks
Continuous speech separation (CSS) is an approach to handling overlapped speech in conversational audio signals.
73 papers · 2 benchmarks
CoVoST2 (Common Voice Speech-To-Text 2)
End-to-end speech-to-text translation (ST) has recently witnessed an increased interest given its system simplicity, lower inference latency and less compounding errors compared to cascaded ST (i.e.
66 papers · 0 benchmarks
The TED-LIUM corpus consists of English-language TED talks.
64 papers · 2 benchmarks
LRS3-TED is a multi-modal dataset for visual and audio-visual speech recognition.
63 papers · 7 benchmarks
WenetSpeech is a multi-domain Mandarin corpus consisting of 10,000+ hours high-quality labeled speech, 2,400+ hours weakly labelled speech, and about 10,000 hours unlabeled speech, with 22,400+ hours in total.
58 papers · 1 benchmark
Europarl-ST is a multilingual Spoken Language Translation corpus containing paired audio-text samples for SLT from and into 9 European languages, for a total of 72 different translation directions.
57 papers · 0 benchmarks
ReVerb Challenge (REverberant Voice Enhancement and Recognition Benchmark)
The REVERB (REverberant Voice Enhancement and Recognition Benchmark) challenge is a benchmark for evaluation of automatic speech recognition techniques.
55 papers · 1 benchmark
VOICES (Voices Obscured In Complex Environmental Settings)
The VOICES corpus is a dataset to promote speech and signal processing research of speech recorded by far-field microphones in noisy room conditions.
48 papers · 0 benchmarks
AliMeeting (Multi-Channel Multi-Party Meeting Transcription Challenge)
AliMeeting corpus consists of 120 hours of recorded Mandarin meeting data, including far-field data collected by 8-channel microphone array as well as near-field data collected by headset microphone.
46 papers · 1 benchmark
CHiME-5 (CHiME Speech Separation and Recognition Challenge)
The CHiME challenge series aims to advance robust automatic speech recognition (ASR) technology by promoting research at the interface of speech and language processing, signal processing , and machine learning.
42 papers · 0 benchmarks
CoVoST is a large-scale multilingual speech-to-text translation corpus.
34 papers · 0 benchmarks
THCHS-30 is a free Chinese speech database THCHS-30 that can be used to build a full-fledged Chinese speech recognition system.
34 papers · 0 benchmarks
2000 HUB5 English Evaluation Transcripts was developed by the Linguistic Data Consortium (LDC) and consists of transcripts of 40 English telephone conversations used in the 2000 HUB5 evaluation sponsored by NIST (National Institute of…
33 papers · 2 benchmarks
NWPU-Crowd consists of 5,109 images, in a total of 2,133,375 annotated heads with points and boxes.
32 papers · 2 benchmarks
TIMIT (TIMIT Acoustic-Phonetic Continuous Speech Corpus)
The TIMIT Acoustic-Phonetic Continuous Speech Corpus is a standard dataset used for evaluation of automatic speech recognition systems.
31 papers · 6 benchmarks
In SpokenSQuAD, the document is in spoken form, the input question is in the form of text and the answer to each question is always a span in the document.
24 papers · 1 benchmark
VocalSound is a free dataset consisting of 21,024 crowdsourced recordings of laughter, sighs, coughs, throat clearing, sneezes, and sniffs from 3,365 unique subjects.
24 papers · 1 benchmark
The Easy Communications (EasyCom) dataset is a world-first dataset designed to help mitigate the cocktail party effect from an augmented-reality (AR) -motivated multi-sensor egocentric world view.
22 papers · 4 benchmarks
SLUE (Spoken Language Understanding Evaluation)
Spoken Language Understanding Evaluation (SLUE) is a suite of benchmark tasks for spoken language understanding evaluation.
22 papers · 3 benchmarks
Taskmaster-1 is a dialog dataset consisting of 13,215 task-based dialogs in English, including 5,507 spoken and 7,708 written dialogs created with two distinct procedures.
19 papers · 0 benchmarks
SPGISpeech (pronounced “speegie-speech”) is a large-scale transcription dataset, freely available for academic research.
16 papers · 1 benchmark
MaSS (Multilingual corpus of Sentence-aligned Spoken utterances) is an extension of the CMU Wilderness Multilingual Speech Dataset, a speech dataset based on recorded readings of the New Testament.
15 papers · 3 benchmarks
SMS-WSJ (Spatialized Multi-Speaker Wall Street Journal)
Spatialized Multi-Speaker Wall Street Journal (SMS-WSJ) consists of artificially mixed speech taken from the WSJ database, but unlike earlier databases this one considers all WSJ0+1 utterances and takes care of strictly separating the…
14 papers · 0 benchmarks
The MagicData-RAMC corpus contains 180 hours of conversational speech data recorded from native speakers of Mandarin Chinese over mobile phones with a sampling rate of 16 kHz.
12 papers · 0 benchmarks
The CALLHOME English Corpus is a collection of unscripted telephone conversations between native speakers of English.
11 papers · 7 benchmarks
Overall duration per microphone: about 36 hours (31 hrs train / 2.5 hrs dev / 2.5 hrs test) Count of microphones: 3 (Microsoft Kinect, Yamaha, Samson) Count of wave-files per microphone: about 14500 Overall count of participations: 180…
11 papers · 1 benchmark
SpeakingFaces is a publicly-available large-scale dataset developed to support multimodal machine learning research in contexts that utilize a combination of thermal, visual, and audio data streams; examples include human-computer…
10 papers · 0 benchmarks
SPEECH-COCO contains speech captions that are generated using text-to-speech (TTS) synthesis resulting in 616,767 spoken captions (more than 600h) paired with images.
9 papers · 0 benchmarks
ADIMA is a novel, linguistically diverse, ethically sourced, expert annotated and well-balanced multilingual profanity detection audio dataset comprising of 11,775 audio samples in 10 Indic languages spanning 65 hours and spoken by 6,446…
8 papers · 0 benchmarks
Europarl-ASR (EN) is a 1300-hour English-language speech and text corpus of parliamentary debates for (streaming) Automatic Speech Recognition training and benchmarking, speech data filtering and speech data verbatimization, based on…
8 papers · 2 benchmarks
JSUT Corpus is a free large-scale speech corpus that can be shared between academic institutions and commercial companies has an important role.
8 papers · 0 benchmarks
The Atari Grand Challenge dataset is a large dataset of human Atari 2600 replays.
7 papers · 0 benchmarks
VIVOS (VIVOS Corpus)
VIVOS is a free Vietnamese speech corpus consisting of 15 hours of recording speech prepared for Automatic Speech Recognition task.
7 papers · 1 benchmark
EmoDB Dataset (Berlin Database of Emotional Speech)
The EMODB database is the freely available German emotional database.
6 papers · 1 benchmark

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.