Home › Datasets › task › Text-To-Speech Synthesis

Text-To-Speech Synthesis datasets

archive 2025-07-28

18 datasets carry the task tag "Text-To-Speech Synthesis" (the task itself: Text-To-Speech Synthesis), ordered by the archive's paper count. Page 1 of 1: 18 shown of 18. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Text-To-Speech Synthesis datasets 1–18 of 18

LJSpeech (The LJ Speech Dataset)
This is a public domain speech dataset consisting of 13,100 short audio clips of a single speaker reading passages from 7 non-fiction books.
323 papers · 2 benchmarks
LibriTTS is a multi-speaker English corpus of approximately 585 hours of read English speech at 24kHz sampling rate, prepared by Heiga Zen with the assistance of Google Speech and Google Brain team members.
257 papers · 1 benchmark
AISHELL-3 is a large-scale and high-fidelity multi-speaker Mandarin speech corpus which could be used to train multi-speaker Text-to-Speech (TTS) systems.
38 papers · 0 benchmarks
CVSS is a massively multilingual-to-English speech to speech translation (S2ST) corpus, covering sentence-level parallel S2ST pairs from 21 languages into English.
26 papers · 1 benchmark
20000 utterances
15 papers · 1 benchmark
SOMOS (The Samsung Open MOS Dataset for the Evaluation of Neural Text-to-Speech Synthesis)
The SOMOS dataset is a large-scale mean opinion scores (MOS) dataset consisting of solely neural text-to-speech (TTS) samples.
11 papers · 0 benchmarks
A large-scale corpus of Gulf Arabic consisting of 110 million words from 1,200 forum novels.
7 papers · 0 benchmarks
KazakhTTS is an open-source speech synthesis dataset for Kazakh, a low-resource language spoken by over 13 million people worldwide.
4 papers · 0 benchmarks
SpeechInstruct is a large-scale cross-modal speech instruction dataset.
4 papers · 0 benchmarks
EMOVIE is a Mandarin emotion speech dataset including 9,724 samples with audio files and its emotion human-labeled annotation.
3 papers · 0 benchmarks
HUI speech corpus (Hof University iisys speech dataset)
The data set contains several speakers.
3 papers · 2 benchmarks
OpenSLR (Open Speech and Language Resources)
OpenSLR is a repository of open speech and language resources, including large-scale transcribed audio corpora and related software.
2 papers · 1 benchmark
RyanSpeech is a speech corpus for research on automated text-to-speech (TTS) systems.
2 papers · 0 benchmarks
Trinity Gesture Dataset includes 23 takes, totalling 244 minutes of motion capture and audio of a male native English speaker producing spontaneous speech on different topics.
2 papers · 2 benchmarks
A Brazilian Portuguese TTS dataset featuring a female voice recorded with high quality in a controlled environment, with neutral emotion and more than 20 hours of recordings.
1 paper · 0 benchmarks
A database containing high sampling rate recordings of a single speaker reading sentences in Brazilian Portuguese with neutral voice, along with the corresponding text corpus.
1 paper · 0 benchmarks
IMaSC (ICFOSS Malayalam Speech Corpus)
IMaSC is a Malayalam text and speech corpus made available by ICFOSS for the purpose of developing speech technology for Malayalam, particularly text-to-speech.
1 paper · 0 benchmarks
Thorsten-Voice (Thorsten-21.02-neutral) is a neutrally spoken voice dataset recorded by Thorsten Müller, audio optimized by Dominik Kreutz and licenced under CC0 to provide it for anybody without any financial or licence struggle.
1 paper · 1 benchmark

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.