Home › Datasets › modality › Speech
Speech datasets
archive 2025-07-28
197 datasets carry the modality tag "Speech", ordered by the archive's paper count. Page 4 of 5: 48 shown of 197. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Speech datasets 145–192 of 197
The Jejueo Single Speaker Speech (JSS) dataset consists of 10k high-quality audio files recorded by a native Jejueo speaker and a transcript file.
1 paper · 0 benchmarks
Kinect-WSJ is a multichannel, multispeaker, reverberated, noisy dataset which extends the WSJ0-2mix singlechannel, non-reverberated, noiseless dataset to the strong reverberation and noise conditions and the Kinect-like microphone array…
1 paper · 0 benchmarks
The Kite database is a multi-modal dataset for the control of unmanned aerial vehicles (UAVs).
1 paper · 0 benchmarks
LibriS2S is a Speech to Speech Translation (S2ST) dataset build further upon existing resources.
1 paper · 0 benchmarks
Here we release the dataset (MultiChannelGrid, abbreviated as MCGrid) used in our paper LIMUSE: LIGHTWEIGHT MULTI-MODAL SPEAKER EXTRACTION](https://arxiv.org/abs/2111.04063)).
1 paper · 0 benchmarks
MSNER (Multilingual Spoken Named Entity Recognition)
This dataset contains named entities annotations for European Parliament recordings in Dutch, French, German and Spanish.
1 paper · 0 benchmarks
MediBeng (Synthetic Code-Switched Bengali-English Speech Conversations for Healthcare Applications)
MediBeng Dataset The MediBeng dataset contains synthetic code-switched dialogues in Bengali and English for training models in speech recognition (ASR), text-to-speech (TTS), and machine translation in clinical settings.
1 paper · 1 benchmark
We announce the release of a new multilingual speaker dataset called NITK-IISc Multilingual Multi-accent Speaker Profiling(NISP) dataset.
1 paper · 0 benchmarks
The NISQA Corpus includes more than 14,000 speech samples with simulated (e.g.
1 paper · 0 benchmarks
Parkinson Speech Dataset is an audio dataset consisting of recordings of 20 Parkinson's Disease (PD) patients and 20 healthy subjects.
1 paper · 0 benchmarks
Data collection was conducted by asking some adults from social media and some students from an elementary school to participate in our experiment.
1 paper · 0 benchmarks
Quechua Collao corpus for automatic emotion recognition in speech.
1 paper · 1 benchmark
Speech Recognition Dataset for Oromo Language.
1 paper · 1 benchmark
Facial electromyography recordings during both silent and vocalized speech.
1 paper · 0 benchmarks
Dataset Summary Speech Brown is a comprehensive, synthetic, and diverse paired speech-text dataset in 15 categories, covering a wide range of topics from fiction to religion.
1 paper · 0 benchmarks
Spot the Difference Corpus is a corpus of task-oriented spontaneous dialogues which contains 54 interactions between pairs of subjects interacting to find differences in two very similar scenes.
1 paper · 0 benchmarks
TAT (Taiwanese Across Taiwan)
Taiwanese Across Taiwan (TAT) corpus is a Large-Scale database of Native Taiwanese Article/Reading Speech collected across Taiwan.
1 paper · 1 benchmark
The TED VCR Video Retrieval Dataset is a multimodal collection derived from publicly available TED Talks.
1 paper · 0 benchmarks
Dubbed series are gaining a lot of popularity in recent years with strong support from major media service providers.
1 paper · 0 benchmarks
100 samples each of synthetic speech generated by 9 moderns TTS systems.
1 paper · 0 benchmarks
The SWC is a corpus of aligned Spoken Wikipedia articles from the English, German, and Dutch Wikipedia.
1 paper · 1 benchmark
The United-Syn-Med dataset is a specialized medical speech dataset designed to evaluate and improve Automatic Speech Recognition (ASR) systems within the healthcare domain.
1 paper · 0 benchmarks
Uses same clean speech as VoiceBank+Demand but more noise types.
1 paper · 1 benchmark
VESUS (Varied Emotion in Syntactically Uniform Speech)
The Varied Emotion in Syntactically Uniform Speech (VESUS) repository is a lexically controlled database collected by the NSA lab.
1 paper · 0 benchmarks
VedantaNY-10M is a curated dataset of over 750 hours of transcripts from public discourses on the Indian philosophy of Advaita Vedanta.
1 paper · 0 benchmarks
This is the forehead accelerometer variant of the VibraVox dataset.
1 paper · 3 benchmarks
This is the reference headset microphone variant of the VibraVox dataset.
1 paper · 2 benchmarks
This is the in-ear rigid earpiece-embedded microphone variant of the VibraVox dataset.
1 paper · 3 benchmarks
This is the in-ear comply foam-embedded microphone variant of the VibraVox dataset.
1 paper · 3 benchmarks
This is the temple vibration pickup variant of the VibraVox dataset.
1 paper · 3 benchmarks
This is the throat microphone (laryngophone) variant of the VibraVox dataset.
1 paper · 3 benchmarks
Voice Navigation is a large-scale dataset of Chinese speech for slot filling, containing more than 830,000 samples.
1 paper · 0 benchmarks
A dataset for voice and 3D face structure study.
1 paper · 1 benchmark
This Sanskrit speech corpus has more than 78 hours of audio data and contains recordings of 45,953 sentences with a sampling rate of 22KHz.
1 paper · 0 benchmarks
WHAMRext is an extension to the WHAMR corpus with larger RT60 values (between 1s and 3s)
1 paper · 1 benchmark
The Watch Your Mouth dataset is a custom silent speech dataset consisting of depth-only recordings of users silently mouthing full English sentences, captured using consumer-grade depth cameras such as the iPhone TrueDepth sensor.
1 paper · 0 benchmarks
The dataset is a private dataset collected for automatic analysis of psychological distress.
1 paper · 1 benchmark
inaGVAD (InaGVAD : a Challenging French TV and Radio Corpus annotated for Voice Activity Detection and Speaker Gender Segmentation)
InaGVAD is a Voice Activity Detection (VAD) and Speaker Gender Segmentation (SGS) dataset designed for representing the acoustic diversity of French TV and Radio programs.
1 paper · 0 benchmarks
A modification on the ShEMO dataset with help of an Automatic Speech Recognition (ASR) system.
1 paper · 0 benchmarks
Dataset based on Twitter usernames of American politicians.
1 paper · 0 benchmarks
The Arabic Speech Corpus (1.5 GB) is a Modern Standard Arabic (MSA) speech corpus for speech synthesis.
0 papers · 0 benchmarks
BABEL is a multilingual corpus of conversational telephone speech from IARPA, which includes Asian and African languages.
0 papers · 0 benchmarks
The CMU Wilderness Multilingual Speech Dataset is a dataset of over 700 different languages providing audio, aligned text and word pronunciations.
0 papers · 0 benchmarks
Deeply Korean read speech corpus contains pairs of Korean speakers reading a script with 3 distinct text sentiments (negative, neutral, positive), with 3 distinct voice sentiments (negative, neutral, positive), are recorded.
0 papers · 0 benchmarks
Deeply Parent-Child Vocal Interaction contains the interaction of 24 pairs of parent and child(total 48 speakers), such as reading fairy tales, singing children’s songs, conversing, and others, is recorded.
0 papers · 0 benchmarks
Deeply vocal characterizer is a human nonverbal vocalization dataset.
0 papers · 0 benchmarks
FluencyBank is a shared database for the study of fluency development.
0 papers · 0 benchmarks
JTES (Japanese Twitter-based Emotional Speech)
We designed an emotional speech database that can be used for emotion recognition as well as recognition and synthesis of speech with various emotions.
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.