Home › Datasets › task › Automatic Speech Recognition
Automatic Speech Recognition datasets
archive 2025-07-28
22 datasets carry the task tag "Automatic Speech Recognition" (the task itself: Automatic Speech Recognition), ordered by the archive's paper count. Page 1 of 1: 22 shown of 22. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Automatic Speech Recognition datasets 1–22 of 22
The MNIST database (Modified National Institute of Standards and Technology database) is a large collection of handwritten digits.
7,651 papers · 44 benchmarks
The LibriSpeech corpus is a collection of approximately 1,000 hours of audiobooks that are a part of the LibriVox project.
2,361 papers · 4 benchmarks
Common Voice is an audio dataset that consists of a unique MP3 and corresponding text file.
449 papers · 143 benchmarks
This is a public domain speech dataset consisting of 13,100 short audio clips of a single speaker reading passages from 7 non-fiction books.
323 papers · 2 benchmarks
FLEURS (Few-shot Learning Evaluation of Universal Representations of Speech)
We introduce FLEURS, the Few-shot Learning Evaluation of Universal Representations of Speech benchmark.
141 papers · 1 benchmark
VoxPopuli is a large-scale multilingual corpus providing 100K hours of unlabelled speech data in 23 languages.
113 papers · 1 benchmark
GigaSpeech, an evolving, multi-domain English speech recognition corpus with 10,000 hours of high quality labeled audio suitable for supervised training, and 40,000 hours of total audio suitable for semi-supervised and unsupervised…
87 papers · 3 benchmarks
Multilingual LibriSpeech is a large multilingual corpus suitable for speech research.
77 papers · 2 benchmarks
The TED-LIUM corpus consists of English-language TED talks.
64 papers · 2 benchmarks
SPGISpeech (pronounced “speegie-speech”) is a large-scale transcription dataset, freely available for academic research.
16 papers · 1 benchmark
The CALLHOME English Corpus is a collection of unscripted telephone conversations between native speakers of English.
11 papers · 7 benchmarks
Earnings-22 is a practical benchmark designed to evaluate automatic speech recognition (ASR) systems' performance on real-world, accented audio.
10 papers · 0 benchmarks
ITALIC: An ITALian Intent Classification Dataset ITALIC is an intent classification dataset for the Italian language, which is the first of its kind.
4 papers · 0 benchmarks
Multilingual automatic speech recognition (ASR) in the medical domain serves as a foundational task for various downstream applications such as speech translation, spoken language understanding, and voice-activated assistants.
3 papers · 0 benchmarks
Jam-ALT (JamALT: A Formatting-Aware Lyrics Transcription Benchmark)
JamALT is a revision of the JamendoLyrics dataset (80 songs in 4 languages), adapted for use as an automatic lyrics transcription (ALT) benchmark.
2 papers · 5 benchmarks
MSC is a dataset for Macro-Management in StarCraft 2 based on the platfrom SC2LE.
2 papers · 0 benchmarks
NPSC (Norwegian Parliamentary Speech Corpus)
The Norwegian Parliamentary Speech Corpus (NPSC) is a speech corpus made by the Norwegian Language Bank at the National Library of Norway in 2019-2021.
2 papers · 0 benchmarks
OpenSLR (Open Speech and Language Resources)
OpenSLR is a repository of open speech and language resources, including large-scale transcribed audio corpora and related software.
2 papers · 1 benchmark
BERSt (Basic Emotion Random phrase Shouts)
BERSt Dataset We release the BERSt Dataset for various speech recognition tasks including Automatic Speech Recognition (ASR) and Speech Emotion Recogniton (SER) Overview 4526 single phrase recordings (~3.75h) 98 professional actors 19…
1 paper · 1 benchmark
MASC (Manually Annotated Sub-Corpus)
The Manually Annotated Sub-Corpus (MASC) consists of approximately 500,000 words of contemporary American English written and spoken data drawn from the Open American National Corpus (OANC).
1 paper · 0 benchmarks
MediBeng (Synthetic Code-Switched Bengali-English Speech Conversations for Healthcare Applications)
MediBeng Dataset The MediBeng dataset contains synthetic code-switched dialogues in Bengali and English for training models in speech recognition (ASR), text-to-speech (TTS), and machine translation in clinical settings.
1 paper · 1 benchmark
Video Dataset (Storytelling Video Dataset (Russian, Emotion, Gesture, Speech))
The Storytelling Video Dataset is a high-quality, human-reviewed multimodal dataset featuring over 700 full-body video recordings of native Russian speakers.
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.