Home › Datasets › modality › Speech
Speech datasets
archive 2025-07-28
197 datasets carry the modality tag "Speech", ordered by the archive's paper count. Page 2 of 5: 48 shown of 197. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Speech datasets 49–96 of 197
The MagicData-RAMC corpus contains 180 hours of conversational speech data recorded from native speakers of Mandarin Chinese over mobile phones with a sampling rate of 16 kHz.
12 papers · 0 benchmarks
BSTC (Baidu Speech Translation Corpus)
BSTC (Baidu Speech Translation Corpus) is a large-scale dataset for automatic simultaneous interpretation.
11 papers · 0 benchmarks
The EARS-WHAM dataset mixes speech from the EARS dataset with real noise recordings from the WHAM!
11 papers · 1 benchmark
SOMOS (The Samsung Open MOS Dataset for the Evaluation of Neural Text-to-Speech Synthesis)
The SOMOS dataset is a large-scale mean opinion scores (MOS) dataset consisting of solely neural text-to-speech (TTS) samples.
11 papers · 0 benchmarks
Overall duration per microphone: about 36 hours (31 hrs train / 2.5 hrs dev / 2.5 hrs test) Count of microphones: 3 (Microsoft Kinect, Yamaha, Samson) Count of wave-files per microphone: about 14500 Overall count of participations: 180…
11 papers · 1 benchmark
VoxForge is an open speech dataset that was set up to collect transcribed speech for use with Free and Open Source Speech Recognition Engines (on Linux, Windows and Mac).
11 papers · 9 benchmarks
UGIF is a multi-lingual, multi-modal UI grounded dataset for step-by-step task completion on the smartphone.
10 papers · 0 benchmarks
Speech encompasses a wealth of information, including but not limited to content, paralinguistic, and environmental information.
9 papers · 0 benchmarks
SPEECH-COCO contains speech captions that are generated using text-to-speech (TTS) synthesis resulting in 616,767 spoken captions (more than 600h) paired with images.
9 papers · 0 benchmarks
SwissDial is an annotated parallel corpus of spoken Swiss German across 8 major dialects, plus a Standard German reference.
9 papers · 0 benchmarks
ADIMA is a novel, linguistically diverse, ethically sourced, expert annotated and well-balanced multilingual profanity detection audio dataset comprising of 11,775 audio samples in 10 Indic languages spanning 65 hours and spoken by 6,446…
8 papers · 0 benchmarks
Emilia Dataset (An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation)
Recent advancements in speech generation models have been significantly driven by the use of large-scale training data.
8 papers · 0 benchmarks
Europarl-ASR (EN) is a 1300-hour English-language speech and text corpus of parliamentary debates for (streaming) Automatic Speech Recognition training and benchmarking, speech data filtering and speech data verbatimization, based on…
8 papers · 2 benchmarks
GigaST is a large-scale pseudo speech translation (ST) corpus.
8 papers · 0 benchmarks
MRDA (ICSI Meeting Recorder Dialog Act Corpus)
The MRDA corpus consists of about 75 hours of speech from 75 naturally-occurring meetings among 53 speakers.
8 papers · 1 benchmark
EasyCall is a new dysarthric speech command dataset in Italian.
7 papers · 0 benchmarks
SingFake (SingFake: Singing Voice Deepfake Detection)
The rise of singing voice synthesis presents critical challenges to artists and industry stakeholders over unauthorized voice usage.
7 papers · 0 benchmarks
The People's Speech is a free-to-download 30,000-hour and growing supervised conversational English speech recognition dataset licensed for academic and commercial usage under CC-BY-SA (with a CC-BY subset).
7 papers · 0 benchmarks
VoicePrivacy 2020 is a dataset for developing anonymization solutions for speech technology.
7 papers · 0 benchmarks
ASCEND (A Spontaneous Chinese-English Dataset) introduces a high-quality resource of spontaneous multi-turn conversational dialogue Chinese code-switching corpus collected in Hong Kong.
6 papers · 0 benchmarks
Common Phone is a gender-balanced, multilingual corpus recorded from more than 76.000 contributors via Mozilla's Common Voice project.
6 papers · 0 benchmarks
FMFCC-A is a large publicly-available Mandarin dataset for synthetic speech detection, which contains 40,000 synthesized Mandarin utterances that generated by 11 Mandarin TTS systems and two Mandarin VC systems, and 10,000 genuine Mandarin…
6 papers · 0 benchmarks
GTSinger (GTSinger: A Global Multi-Technique Singing Corpus with Realistic Music Scores for All Singing Tasks)
The scarcity of high-quality and multi-task singing datasets significantly hinders the development of diverse controllable and personalized singing tasks, as existing singing datasets suffer from low quality, limited diversity of languages…
6 papers · 0 benchmarks
A large-scale (105K conversations) media dialog dataset collected from news interview transcripts.
6 papers · 0 benchmarks
Libri-adhoc40 is a synchronized speech corpus which collects the replayed Librispeech data from loudspeakers by ad-hoc microphone arrays of 40 strongly synchronized distributed nodes in a real office environment.
6 papers · 0 benchmarks
Open Images is a computer vision dataset covering ~9 million images with labels spanning thousands of object categories.
6 papers · 0 benchmarks
SpeechMatrix is a large-scale multilingual corpus of speech-to-speech translations mined from real speech of European Parliament recordings.
6 papers · 0 benchmarks
CSRC (Children Speech Recognition Challenge)
CSRC is a collection of data for Children Speech Recognition.
5 papers · 0 benchmarks
ClovaCall is a new large-scale Korean call-based speech corpus under a goal-oriented dialog scenario from more than 11,000 people.
5 papers · 0 benchmarks
DeToxy (DeToxy: A Large-Scale Multimodal Dataset for Toxicity Classification in Spoken Utterances)
DeToxy is a publicly available toxicity annotated dataset for the English language.
5 papers · 0 benchmarks
Libri-Adapt aims to support unsupervised domain adaptation research on speech recognition models.
5 papers · 0 benchmarks
MediaSpeech is a media speech dataset (you might have guessed this) built with the purpose of testing Automated Speech Recognition (ASR) systems performance.
5 papers · 1 benchmark
NISP (NITK-IISc Multilingual Multi-accent Speaker Profiling)
This dataset contains speech recordings along with speaker physical parameters (height, weight, shoulder size, age ) as well as regional information and linguistic information.
5 papers · 0 benchmarks
The OLR 2021 dataset contains the data for the Oriental Language Recognition (OLR) 2021 Challenge, which intends to improve the performance of language recognition systems and speech recognition systems within multilingual scenarios.
5 papers · 0 benchmarks
The PodcastFillers dataset consists of 199 full-length podcast episodes in English with manually annotated filler words and automatically generated transcripts.
5 papers · 1 benchmark
RealMAN (A Real-Recorded and Annotated Microphone Array Dataset for Dynamic Speech Enhancement and Localization)
The Audio Signal and Information Processing Lab at Westlake University, in collaboration with AISHELL, has released the Real-recorded and annotated Microphone Array speech&Noise (RealMAN) dataset, which provides annotated multi-channel…
5 papers · 2 benchmarks
This is a 16.2-million frame (50-hour) multimodal dataset of two-person face-to-face spontaneous conversations.
5 papers · 0 benchmarks
Timers and Such is an open source dataset of spoken English commands for common voice control use cases involving numbers.
5 papers · 1 benchmark
A large-scale corpus for phonetic typology, with aligned segments and estimated phoneme-level labels in 690 readings spanning 635 languages, along with acoustic-phonetic measures of vowels and sibilants.
5 papers · 0 benchmarks
DISRPT2019 (DISRPT2019 shared task on Discourse Unit Segmentation and Connective Detection)
The DISRPT 2019 workshop introduces the first iteration of a cross-formalism shared task on discourse unit segmentation.
4 papers · 0 benchmarks
EdAcc (Edinburgh International Accents of English Corpus)
The Edinburgh International Accents of English Corpus (EdAcc) is a new automatic speech recognition (ASR) dataset composed of 40 hours of English dyadic conversations between speakers with a diverse set of accents.
4 papers · 0 benchmarks
ExVo2022 (ICML ExVo 2022 Workshop & Competition Data)
Baseline code for the three tracks of ExVo 2022 competition.
4 papers · 0 benchmarks
KazakhTTS is an open-source speech synthesis dataset for Kazakh, a low-resource language spoken by over 13 million people worldwide.
4 papers · 0 benchmarks
Kosp2e (read as kospi'), is a corpus that allows Korean speech to be translated into English text in an end-to-end manner
4 papers · 0 benchmarks
Real-M is a crowd-sourced speech-separation corpus of real-life mixtures.
4 papers · 0 benchmarks
RTASC (ROBIN Technical Acquisition Speech Corpus)
The ROBIN Technical Acquisition Speech Corpus (ROBINTASC) was developed within the ROBIN project.
4 papers · 0 benchmarks
We introduce a new database of voice recordings with the goal of supporting research on vulnerabilities and protection of voice-controlled systems.
4 papers · 0 benchmarks
SpeechInstruct is a large-scale cross-modal speech instruction dataset.
4 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.