Home › Datasets › task › Speech Recognition

Speech Recognition datasets

archive 2025-07-28

96 datasets carry the task tag "Speech Recognition" (the task itself: Speech Recognition), ordered by the archive's paper count. Page 2 of 2: 48 shown of 96. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Speech Recognition datasets 49–96 of 96

Open Images is a computer vision dataset covering ~9 million images with labels spanning thousands of object categories.
6 papers · 0 benchmarks
The ADL Piano MIDI is a dataset of 11,086 piano pieces from different genres.
5 papers · 0 benchmarks
AccentDB is a database that contains samples of 4 Indian-English accents, and a compilation of samples from 4 native-English, and a metropolitan Indian-English accent.
5 papers · 0 benchmarks
Artie Bias Corpus is an open dataset for detecting demographic bias in speech applications.
5 papers · 0 benchmarks
CSRC (Children Speech Recognition Challenge)
CSRC is a collection of data for Children Speech Recognition.
5 papers · 0 benchmarks
ClovaCall is a new large-scale Korean call-based speech corpus under a goal-oriented dialog scenario from more than 11,000 people.
5 papers · 0 benchmarks
Libri-Adapt aims to support unsupervised domain adaptation research on speech recognition models.
5 papers · 0 benchmarks
MediaSpeech is a media speech dataset (you might have guessed this) built with the purpose of testing Automated Speech Recognition (ASR) systems performance.
5 papers · 1 benchmark
The OLR 2021 dataset contains the data for the Oriental Language Recognition (OLR) 2021 Challenge, which intends to improve the performance of language recognition systems and speech recognition systems within multilingual scenarios.
5 papers · 0 benchmarks
Timers and Such is an open source dataset of spoken English commands for common voice control use cases involving numbers.
5 papers · 1 benchmark
word2word contains easy-to-use word translations for 3,564 language pairs.
5 papers · 0 benchmarks
RTASC (ROBIN Technical Acquisition Speech Corpus)
The ROBIN Technical Acquisition Speech Corpus (ROBINTASC) was developed within the ROBIN project.
4 papers · 0 benchmarks
VietMed (VietMed: A Dataset and Benchmark for Automatic Speech Recognition of Vietnamese in the Medical Domain)
We introduced a Vietnamese speech recognition dataset in the medical domain comprising 16h of labeled medical speech, 1000h of unlabeled medical speech and 1200h of unlabeled general-domain speech.
4 papers · 2 benchmarks
ArzEn (Corpus of Egyptian Arabic-English Code-switching)
Corpus of Egyptian Arabic-English Code-switching (ArzEn) is a spontaneous conversational speech corpus, obtained through informal interviews held at the German University in Cairo.
3 papers · 0 benchmarks
This noisy speech test set is created from the Google Speech Commands v2 [1] and the Musan dataset[2].
3 papers · 1 benchmark
LibriVoxDeEn is a corpus of sentence-aligned triples of German audio, German text, and English translation, based on German audiobooks.
3 papers · 0 benchmarks
NusaCrowd is a collaborative initiative to collect and unite existing resources for Indonesian languages, including opening access to previously non-public resources.
3 papers · 0 benchmarks
ODSQA (Open-Domain Spoken Question Answering)
The ODSQA dataset is a spoken dataset for question answering in Chinese.
3 papers · 0 benchmarks
RESD (Russian Emotional Speech Dialogs with annotated text)
Russian dataset of emotional speech dialogues.
3 papers · 1 benchmark
TaL Corpus (The Tongue and Lips Corpus)
The Tongue and Lips (TaL) corpus is a multi-speaker corpus of ultrasound images of the tongue and video images of lips.
3 papers · 0 benchmarks
This work introduces Zambezi Voice, an open-source multilingual speech resource for Zambian languages.
3 papers · 0 benchmarks
AV Digits Database is an audiovisual database which contains normal, whispered and silent speech.
2 papers · 0 benchmarks
BD-4SK-ASR (Basic Dataset for Sorani Kurdish Automatic Speech Recognition)
The Basic Dataset for Sorani Kurdish Automatic Speech Recognition (BD-4SK-ASR) is a dataset for automatic speech recognition for Sorani Kurdish.
2 papers · 0 benchmarks
ESB (End-to-End Speech Benchmark)
ESB is a benchmark for evaluating the performance of a single automatic speech recognition (ASR) system across a broad set of speech datasets.
2 papers · 0 benchmarks
Golos is a Russian speech dataset suitable for speech research.
2 papers · 0 benchmarks
Jam-ALT (JamALT: A Formatting-Aware Lyrics Transcription Benchmark)
JamALT is a revision of the JamendoLyrics dataset (80 songs in 4 languages), adapted for use as an automatic lyrics transcription (ALT) benchmark.
2 papers · 5 benchmarks
MASRI-HEADSET is a corpus that was developed by the MASRI project at the University of Malta.
2 papers · 0 benchmarks
NeuroVoz (NeuroVoz: a Castillian Spanish corpus of parkinsonian speech)
The NeuroVoz dataset emerges as a pioneering resource in the field of computational linguistics and biomedical research, specifically designed to enhance the diagnosis and understanding of Parkinson's Disease (PD) through speech analysis.
2 papers · 0 benchmarks
OpenSLR (Open Speech and Language Resources)
OpenSLR is a repository of open speech and language resources, including large-scale transcribed audio corpora and related software.
2 papers · 1 benchmark
Tilde MODEL Corpus (Tilde Multilingual Open Data for European Languages)
Tilde MODEL Corpus is a multilingual corpora for European languages – particularly focused on the smaller languages.
2 papers · 0 benchmarks
This dataset is designed to help train simple machine learning models that serve educational and research purposes in the speech recognition domain, mainly for keyword spotting tasks.
1 paper · 0 benchmarks
BERSt (Basic Emotion Random phrase Shouts)
BERSt Dataset We release the BERSt Dataset for various speech recognition tasks including Automatic Speech Recognition (ASR) and Speech Emotion Recogniton (SER) Overview 4526 single phrase recordings (~3.75h) 98 professional actors 19…
1 paper · 1 benchmark
A new large-scale, in-thewild Mandarin dataset, CAS-VSR-S101 with 101.1 hours of data.
1 paper · 3 benchmarks
CUCO Database (A voice and speech corpus of patients who underwent upper airway surgery in pre-and post-operative states)
Many research articles have explored the impact of surgical interventions on voice and speech evaluations, but advances are limited by the lack of publicly accessible datasets.
1 paper · 0 benchmarks
CrowdSpeech is a publicly available large-scale dataset of crowdsourced audio transcriptions.
1 paper · 2 benchmarks
DEEP-VOICE: Real-time Detection of AI-Generated Speech for DeepFake Voice Conversion This dataset contains examples of real human speech, and DeepFake versions of those speeches by using Retrieval-based Voice Conversion.
1 paper · 1 benchmark
EmoSpeech contains keywords with diverse emotions and background sounds, presented to explore new challenges in audio analysis.
1 paper · 0 benchmarks
FT Speech is a speech corpus created from the recorded meetings of the Danish Parliament, otherwise known as the Folketing (FT).
1 paper · 0 benchmarks
Fongbe audio (Fongbe dataset)
Fongbe Data collected by Fréjus A.
1 paper · 1 benchmark
The Kite database is a multi-modal dataset for the control of unmanned aerial vehicles (UAVs).
1 paper · 0 benchmarks
MSNER (Multilingual Spoken Named Entity Recognition)
This dataset contains named entities annotations for European Parliament recordings in Dutch, French, German and Spanish.
1 paper · 0 benchmarks
MediBeng (Synthetic Code-Switched Bengali-English Speech Conversations for Healthcare Applications)
MediBeng Dataset The MediBeng dataset contains synthetic code-switched dialogues in Bengali and English for training models in speech recognition (ASR), text-to-speech (TTS), and machine translation in clinical settings.
1 paper · 1 benchmark
Data collection was conducted by asking some adults from social media and some students from an elementary school to participate in our experiment.
1 paper · 0 benchmarks
RadioTalk is a corpus of speech recognition transcripts sampled from talk radio broadcasts in the United States between October of 2018 and March of 2019.
1 paper · 0 benchmarks
The United-Syn-Med dataset is a specialized medical speech dataset designed to evaluate and improve Automatic Speech Recognition (ASR) systems within the healthcare domain.
1 paper · 0 benchmarks
Vāksañcayaḥ (Sanskrit Speech Corpus by IIT Bombay)
This Sanskrit speech corpus has more than 78 hours of audio data and contains recordings of 45,953 sentences with a sampling rate of 22KHz.
1 paper · 0 benchmarks
i3-video (is-it-instructional-video)
The i3-video dataset contains "is-it-instructional" annotations for 6.4k videos from Youtube-8M.
1 paper · 0 benchmarks
A Chinese Mandarin speech corpus by Beijing DataTang Technology Co., Ltd, containing 200 hours of speech data from 600 speakers.
0 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.