Datasets › LibriTTS

LibriTTS

Introduced in LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech5 Apr 2019 archive 2025-07-28

LibriTTS is a multi-speaker English corpus of approximately 585 hours of read English speech at 24kHz sampling rate, prepared by Heiga Zen with the assistance of Google Speech and Google Brain team members. The LibriTTS corpus is designed for TTS research. It is derived from the original materials (mp3 audio files from LibriVox and text files from Project Gutenberg) of the LibriSpeech corpus. The main differences from the LibriSpeech corpus are listed below:

  • The audio files are at 24kHz sampling rate.
  • The speech is split at sentence breaks.
  • Both original and normalized texts are included.
  • Contextual information (e.g., neighbouring sentences) can be extracted.
  • Utterances with significant background noise are excluded.

Benchmarks archive 2025-07-28

All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Speech Synthesis LibriTTS PeriodWave-Turbo-L PESQ 4.454 Accelerating High-Fidelity Waveform Generation via... sh-lee-prml/periodwave 15 Compare

Papers archive 2025-07-28

11 shown of 11 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 257. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Accelerating High-Fidelity Waveform Generation via Adversarial Flow Matching Optimization 1 1 15 Aug 2024 not harvested
PeriodWave: Multi-Period Flow Matching for High-Fidelity Waveform Generation 1 1 14 Aug 2024 not harvested
RFWave: Multi-band Rectified Flow for Audio Waveform Reconstruction 1 1 8 Mar 2024 ran 1 of 4 samples (3 unverified)
EVA-GAN: Enhanced Various Audio Generation via Scalable Generative Adversarial Networks 1 2 31 Jan 2024 not harvested
BigVSAN: Enhancing GAN-based Neural Vocoders with Slicing Adversarial Network 3 2 6 Sep 2023 not harvested
Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis 4 1 1 Jun 2023 ran 5 of 16 samples (11 unverified)
BigVGAN: A Universal Neural Vocoder with Large-Scale Training 5 3 9 Jun 2022 ran 8 of 17 samples (9 unverified; 1 pointer-only for licence)
HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis 11 1 12 Oct 2020 ran 17 of 25 samples (8 unverified; 3 pointer-only for licence)
Speaker Conditional WaveRNN: Towards Universal Neural Vocoder for Unseen Speaker and Recording Conditions 1 1 9 Aug 2020 ran 1 of 1 samples (0 unverified; 1 pointer-only for licence)
WaveFlow: A Compact Flow-based Model for Raw Audio 4 1 3 Dec 2019 ran 3 of 9 samples (6 unverified)
WaveGlow: A Flow-based Generative Network for Speech Synthesis 2 1 31 Oct 2018 ran 2 of 7 samples (5 unverified)

Dataset loaders archive 2025-07-28

2 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • LibriTTS

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections