Datasets › LJSpeech

LJSpeech (The LJ Speech Dataset)

Introduced in The lj speech dataset archive 2025-07-28

This is a public domain speech dataset consisting of 13,100 short audio clips of a single speaker reading passages from 7 non-fiction books. A transcription is provided for each clip. Clips vary in length from 1 to 10 seconds and have a total length of approximately 24 hours. The texts were published between 1884 and 1964, and are in the public domain. The audio was recorded in 2016-17 by the LibriVox project and is also in the public domain.

Source: The LJ Speech Dataset Image Source: https://keithito.com/LJ-Speech-Dataset/ Audio Source: https://keithito.com/LJ-Speech-Dataset/

Benchmarks archive 2025-07-28

All 2 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Text-To-Speech Synthesis LJSpeech NaturalSpeech Audio Quality MOS 4.56 NaturalSpeech: End-to-End Text to Speech Synthesis with... microsoft/NeuralSpeech +2 16 Compare
Speech Synthesis LJSpeech BDDM vocoder Mean Opinion Score 4.48 BDDM: Bilateral Denoising Diffusion Models for Fast and... tencent-ailab/bddm 4 Compare

Papers archive 2025-07-28

13 shown of 13 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 323. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Matcha-TTS: A fast TTS architecture with conditional flow matching 1 1 6 Sep 2023 ran 2 of 2 samples (0 unverified)
OverFlow: Putting flows on top of neural transducers for better TTS 2 1 13 Nov 2022 not harvested
NaturalSpeech: End-to-End Text to Speech Synthesis with Human-Level Quality 3 3 9 May 2022 ran 0 of 3 samples (3 unverified)
FastDiff: A Fast Conditional Diffusion Model for High-Quality Speech Synthesis 2 2 21 Apr 2022 ran 1 of 4 samples (3 unverified; 4 pointer-only for licence)
BDDM: Bilateral Denoising Diffusion Models for Fast and High-Quality Speech Synthesis 1 1 25 Mar 2022 ran 5 of 6 samples (1 unverified)
Neural HMMs are all you need (for high-quality attention-free TTS) 2 2 30 Aug 2021 not harvested
Grad-TTS: A Diffusion Probabilistic Model for Text-to-Speech 6 1 13 May 2021 ran 9 of 18 samples (9 unverified; 1 pointer-only for licence)
DiffWave: A Versatile Diffusion Model for Audio Synthesis 11 1 21 Sep 2020 ran 20 of 33 samples (13 unverified)
FastSpeech 2: Fast and High-Quality End-to-End Text to Speech 37 1 8 Jun 2020 ran 73 of 119 samples (46 unverified; 33 pointer-only for licence)
Glow-TTS: A Generative Flow for Text-to-Speech via Monotonic Alignment Search 6 1 22 May 2020 ran 3 of 14 samples (11 unverified)
Flowtron: an Autoregressive Flow-based Generative Network for Text-to-Speech Synthesis 3 2 12 May 2020 ran 4 of 19 samples (15 unverified)
FastSpeech: Fast, Robust and Controllable Text to Speech 22 2 22 May 2019 ran 3 of 11 samples (8 unverified; 3 pointer-only for licence)
Neural Speech Synthesis with Transformer Network 6 1 19 Sep 2018 not harvested

Dataset loaders archive 2025-07-28

4 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

Public domain

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • LJSpeech

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections