Home › Datasets › modality › Speech

Speech datasets

archive 2025-07-28

197 datasets carry the modality tag "Speech", ordered by the archive's paper count. Page 3 of 5: 48 shown of 197. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Speech datasets 97–144 of 197

VietMed (VietMed: A Dataset and Benchmark for Automatic Speech Recognition of Vietnamese in the Medical Domain)
We introduced a Vietnamese speech recognition dataset in the medical domain comprising 16h of labeled medical speech, 1000h of unlabeled medical speech and 1200h of unlabeled general-domain speech.
4 papers · 2 benchmarks
The ASR-GLUE benchmark is a collection of 6 different NLU (Natural Language Understanding) tasks for evaluating the performance of models under automatic speech recognition (ASR) error across 3 different levels of background noise and 6…
3 papers · 0 benchmarks
DISRPT2021 (DISRPT2021 shared task on Discourse Unit Segmentation, Connective Detection and Discourse Relation Classification)
The DISRPT 2021 shared task, co-located with CODI 2021 at EMNLP, introduces the second iteration of a cross-formalism shared task on discourse unit segmentation and connective detection, as well as the first iteration of a cross-formalism…
3 papers · 0 benchmarks
EMOVIE is a Mandarin emotion speech dataset including 9,724 samples with audio files and its emotion human-labeled annotation.
3 papers · 0 benchmarks
This noisy speech test set is created from the Google Speech Commands v2 [1] and the Musan dataset[2].
3 papers · 1 benchmark
LibriVoxDeEn is a corpus of sentence-aligned triples of German audio, German text, and English translation, based on German audiobooks.
3 papers · 0 benchmarks
Operating rooms (ORs) are complex, high-stakes environments requiring precise understanding of interactions among medical staff, tools, and equipment for enhancing surgical assistance, situational awareness, and patient safety.
3 papers · 3 benchmarks
NusaCrowd is a collaborative initiative to collect and unite existing resources for Indonesian languages, including opening access to previously non-public resources.
3 papers · 0 benchmarks
RESD (Russian Emotional Speech Dialogs with annotated text)
Russian dataset of emotional speech dialogues.
3 papers · 1 benchmark
SDN (Situated Dialogue Navigation)
Situated Dialogue Navigation (SDN) is a navigation benchmark of 183 trials with a total of 8415 utterances, around 18.7 hours of control streams, and 2.9 hours of trimmed audio.
3 papers · 0 benchmarks
Spoken versions of the Semantic Textual Similarity dataset for testing semantic sentence level embeddings.
3 papers · 0 benchmarks
TaL Corpus (The Tongue and Lips Corpus)
The Tongue and Lips (TaL) corpus is a multi-speaker corpus of ultrasound images of the tongue and video images of lips.
3 papers · 0 benchmarks
AV Digits Database is an audiovisual database which contains normal, whispered and silent speech.
2 papers · 0 benchmarks
BD-4SK-ASR (Basic Dataset for Sorani Kurdish Automatic Speech Recognition)
The Basic Dataset for Sorani Kurdish Automatic Speech Recognition (BD-4SK-ASR) is a dataset for automatic speech recognition for Sorani Kurdish.
2 papers · 0 benchmarks
Cantonese In-car Audio-Visual Speech Recognition (CI-AVSR) is a dataset for in-car command recognition in the Cantonese language with both video and audio data.
2 papers · 0 benchmarks
The EARS-Reverb dataset uses real recorded room impulse responses (RIRs) from multiple public datasets (ACE-Challenge, AIR, ARNI, BRUDEX, dEchorate, DetmoldSRIR, and Palimpsest).
2 papers · 1 benchmark
ESB (End-to-End Speech Benchmark)
ESB is a benchmark for evaluating the performance of a single automatic speech recognition (ASR) system across a broad set of speech datasets.
2 papers · 0 benchmarks
Fingerprint Dataset (Neural Audio Fingerprint Dataset)
This dataset includes all music sources, background noises and impulse-reponses (IR) samples and conversation speech that have been used in the work "Neural Audio Fingerprint for High-specific Audio Retrieval based on Contrastive Learning"…
2 papers · 0 benchmarks
The Flickr 8k Audio Caption Corpus contains 40,000 spoken captions of 8,000 natural images.
2 papers · 0 benchmarks
Golos is a Russian speech dataset suitable for speech research.
2 papers · 0 benchmarks
We release the dataset for non-commercial research.
2 papers · 0 benchmarks
JVS-MuSiC is a Japanese multispeaker singing-voice corpus called "JVS-MuSiC" with the aim to analyze and synthesize a variety of voices.
2 papers · 0 benchmarks
Jam-ALT (JamALT: A Formatting-Aware Lyrics Transcription Benchmark)
JamALT is a revision of the JamendoLyrics dataset (80 songs in 4 languages), adapted for use as an automatic lyrics transcription (ALT) benchmark.
2 papers · 5 benchmarks
LaboroTVSpeech is a large-scale Japanese speech corpus built from broadcast TV recordings and their subtitles.
2 papers · 0 benchmarks
Despite impressive advancements in video understanding, most efforts remain limited to coarse-grained or visual-only video tasks.
2 papers · 0 benchmarks
MASRI-HEADSET is a corpus that was developed by the MASRI project at the University of Malta.
2 papers · 0 benchmarks
MAVS (Multilingual Audio-Visual Smartphone dataset)
MAVS is an audio-visual smartphone dataset captured in five different recent smartphones.
2 papers · 0 benchmarks
MultiSV is a corpus designed for training and evaluating text-independent multi-channel speaker verification systems.
2 papers · 0 benchmarks
NPSC (Norwegian Parliamentary Speech Corpus)
The Norwegian Parliamentary Speech Corpus (NPSC) is a speech corpus made by the Norwegian Language Bank at the National Library of Norway in 2019-2021.
2 papers · 0 benchmarks
NeuroVoz (NeuroVoz: a Castillian Spanish corpus of parkinsonian speech)
The NeuroVoz dataset emerges as a pioneering resource in the field of computational linguistics and biomedical research, specifically designed to enhance the diagnosis and understanding of Parkinson's Disease (PD) through speech analysis.
2 papers · 0 benchmarks
RUSLAN is a Russian spoken language corpus for text-to-speech task.
2 papers · 0 benchmarks
RyanSpeech is a speech corpus for research on automated text-to-speech (TTS) systems.
2 papers · 0 benchmarks
VocBench is a framework that benchmark the performance of state-of-the art neural vocoders.
2 papers · 0 benchmarks
Overview nEMO is a simulated dataset of emotional speech in the Polish language.
2 papers · 0 benchmarks
AVASpeech-SMAD (AVASpeech-SMAD: A Strongly Labelled Speech and Music Activity Detection Dataset with Label Co-Occurrence)
We propose a dataset, AVASpeech-SMAD, to assist speech and music activity detection research.
1 paper · 0 benchmarks
This dataset is designed to help train simple machine learning models that serve educational and research purposes in the speech recognition domain, mainly for keyword spotting tasks.
1 paper · 0 benchmarks
Auto-KWS is a dataset for customized keyword spotting, the task of detecting spoken keywords.
1 paper · 0 benchmarks
A new large-scale, in-thewild Mandarin dataset, CAS-VSR-S101 with 101.1 hours of data.
1 paper · 3 benchmarks
CLIPS (Corpora e Lessici dell'Italiano Parlato e Scritto)
CLIPS, ovvero Corpora e Lessici dell'Italiano Parlato e Scritto, è uno degli otto progetti (Progetto n.
1 paper · 0 benchmarks
CSI is a criminal conversational dataset for speaker identification built from the CSI television show.
1 paper · 0 benchmarks
CUCO Database (A voice and speech corpus of patients who underwent upper airway surgery in pre-and post-operative states)
Many research articles have explored the impact of surgical interventions on voice and speech evaluations, but advances are limited by the lack of publicly accessible datasets.
1 paper · 0 benchmarks
CrowdSpeech is a publicly available large-scale dataset of crowdsourced audio transcriptions.
1 paper · 2 benchmarks
DR-VCTK (Device Recorded VCTK)
This dataset is a new variant of the voice cloning toolkit (VCTK) dataset: device-recorded VCTK (DR-VCTK), where the high-quality speech signals recorded in a semi-anechoic chamber using professional audio devices are played back and…
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
The EVI dataset is a challenging, multilingual spoken-dialogue dataset with 5,506 dialogues in English, Polish, and French.
1 paper · 3 benchmarks
EmoSpeech contains keywords with diverse emotions and background sounds, presented to explore new challenges in audio analysis.
1 paper · 0 benchmarks
FT Speech is a speech corpus created from the recorded meetings of the Danish Parliament, otherwise known as the Folketing (FT).
1 paper · 0 benchmarks
Greek Parliament Proceedings is a curated dataset of the Greek Parliament Proceedings that extends chronologically from 1989 up to 2020.
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.