Home › Datasets › task › Speech Synthesis

Speech Synthesis datasets

archive 2025-07-28

22 datasets carry the task tag "Speech Synthesis" (the task itself: Speech Synthesis), ordered by the archive's paper count. Page 1 of 1: 22 shown of 22. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Speech Synthesis datasets 1–22 of 22

LJSpeech (The LJ Speech Dataset)
This is a public domain speech dataset consisting of 13,100 short audio clips of a single speaker reading passages from 7 non-fiction books.
323 papers · 2 benchmarks
LibriTTS is a multi-speaker English corpus of approximately 585 hours of read English speech at 24kHz sampling rate, prepared by Heiga Zen with the assistance of Google Speech and Google Brain team members.
257 papers · 1 benchmark
THCHS-30 is a free Chinese speech database THCHS-30 that can be used to build a full-fledged Chinese speech recognition system.
34 papers · 0 benchmarks
A collection of single speaker speech datasets for ten languages.
22 papers · 0 benchmarks
PromptSpeech is a dataset that consists of speech and the corresponding prompts.
14 papers · 0 benchmarks
SOMOS (The Samsung Open MOS Dataset for the Evaluation of Neural Text-to-Speech Synthesis)
The SOMOS dataset is a large-scale mean opinion scores (MOS) dataset consisting of solely neural text-to-speech (TTS) samples.
11 papers · 0 benchmarks
JSUT Corpus is a free large-scale speech corpus that can be shared between academic institutions and commercial companies has an important role.
8 papers · 0 benchmarks
A large-scale corpus of Gulf Arabic consisting of 110 million words from 1,200 forum novels.
7 papers · 0 benchmarks
Blizzard Challenge 2013 (Blizzard Challenge 2013 - English language tasks)
The English data for voice building was obtained, prepared and provided the the challenge by Lessac Technologies Inc., having originally came from the publishers Voice Factory International Inc.
5 papers · 1 benchmark
HUI speech corpus (Hof University iisys speech dataset)
The data set contains several speakers.
3 papers · 2 benchmarks
TaL Corpus (The Tongue and Lips Corpus)
The Tongue and Lips (TaL) corpus is a multi-speaker corpus of ultrasound images of the tongue and video images of lips.
3 papers · 0 benchmarks
JIT Dataset (Jejueo Interview Transcripts)
The Jejueo Interview Transcripts (JIT) dataset is a parallel corpus containing 170k+ Jejueo-Korean sentences.
2 papers · 0 benchmarks
JVS-MuSiC is a Japanese multispeaker singing-voice corpus called "JVS-MuSiC" with the aim to analyze and synthesize a variety of voices.
2 papers · 0 benchmarks
RUSLAN is a Russian spoken language corpus for text-to-speech task.
2 papers · 0 benchmarks
The dataset has 10.5 hours from a single speaker.
2 papers · 0 benchmarks
Tilde MODEL Corpus (Tilde Multilingual Open Data for European Languages)
Tilde MODEL Corpus is a multilingual corpora for European languages – particularly focused on the smaller languages.
2 papers · 0 benchmarks
VocBench is a framework that benchmark the performance of state-of-the art neural vocoders.
2 papers · 0 benchmarks
A database containing high sampling rate recordings of a single speaker reading sentences in Brazilian Portuguese with neutral voice, along with the corresponding text corpus.
1 paper · 0 benchmarks
JSS Dataset (Jejueo Single Speaker Speech)
The Jejueo Single Speaker Speech (JSS) dataset consists of 10k high-quality audio files recorded by a native Jejueo speaker and a transcript file.
1 paper · 0 benchmarks
Facial electromyography recordings during both silent and vocalized speech.
1 paper · 0 benchmarks
100 samples each of synthetic speech generated by 9 moderns TTS systems.
1 paper · 0 benchmarks
The United-Syn-Med dataset is a specialized medical speech dataset designed to evaluate and improve Automatic Speech Recognition (ASR) systems within the healthcare domain.
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.