Datasets › LibriSpeech

LibriSpeech

Introduced in Librispeech: An ASR corpus based on public domain audio books1 Jan 2015 archive 2025-07-28

The LibriSpeech corpus is a collection of approximately 1,000 hours of audiobooks that are a part of the LibriVox project. Most of the audiobooks come from the Project Gutenberg. The training data is split into 3 partitions of 100hr, 360hr, and 500hr sets while the dev and test data are split into the ’clean’ and ’other’ categories, respectively, depending upon how well or challenging Automatic Speech Recognition systems would perform against. Each of the dev and test sets is around 5hr in audio length. This corpus also provides the n-gram language models and the corresponding texts excerpted from the Project Gutenberg books, which contain 803M tokens and 977K unique words.

Source: State-of-the-art Speech Recognition using Multi-stream Self-attention with Dilated 1D Convolutions

Benchmarks archive 2025-07-28

All 4 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Speech Recognition LibriSpeech test-clean United Med ASR Word Error Rate (WER) 0.985 High-precision medical speech recognition through... — 64 Compare
Speech Recognition LibriSpeech test-other SAMBA ASR Word Error Rate (WER) 2.48 Samba-ASR: State-Of-The-Art Speech Recognition... — 53 Compare
Voice Conversion LibriSpeech test-clean kNN-VC (prematched HiFiGAN) Character Error Rate (CER) 2.96 Voice Conversion With Just Nearest Neighbors bshall/knn-vc 1 Compare
Speech Recognition Librispeech (other) no rows — — 0 Compare

Papers archive 2025-07-28

30 shown of 53 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 2,361. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Let SSMs be ConvNets: State-space Modeling with Optimal Tensor Contractions 0 1 22 Jan 2025 not harvested
Samba-ASR: State-Of-The-Art Speech Recognition Leveraging Structured State-Space Models 0 2 6 Jan 2025 not harvested
High-precision medical speech recognition through synthetic data and semantic correction: UNITED-MEDASR 0 1 24 Nov 2024 not harvested
CR-CTC: Consistency regularization on CTC for improved speech recognition 1 4 7 Oct 2024 not harvested
FAdam: Adam is a natural gradient optimizer using diagonal empirical Fisher information 1 2 21 May 2024 not harvested
Graph Convolutions Enrich the Self-Attention in Transformers! 1 2 7 Dec 2023 ran 19 of 29 samples (10 unverified)
Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models 2 2 14 Nov 2023 ran 5 of 7 samples (2 unverified; 7 pointer-only for licence)
Zipformer: A faster and better encoder for automatic speech recognition 1 2 17 Oct 2023 not harvested
Voice Conversion With Just Nearest Neighbors 1 1 30 May 2023 not harvested
Multi-Head State Space Model for Speech Recognition 0 1 21 May 2023 not harvested
Fast Conformer with Linearly Scalable Attention for Efficient Speech Recognition 0 1 8 May 2023 not harvested
MT4SSL: Boosting Self-Supervised Speech Representation Learning by Integrating Multiple Targets 1 2 14 Nov 2022 not harvested
E-Branchformer: Branchformer with Enhanced merging for speech recognition 1 2 30 Sep 2022 not harvested
Squeezeformer: An Efficient Transformer for Automatic Speech Recognition 4 2 2 Jun 2022 ran 31 of 49 samples (18 unverified)
data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language 12 1 7 Feb 2022 ran 0 of 6 samples (6 unverified)
WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing 9 2 26 Oct 2021 not harvested
W2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre-Training 4 2 7 Aug 2021 not harvested
Amortized Neural Networks for Low-Latency Speech Recognition 0 1 3 Aug 2021 not harvested
Relaxed Attention: A Simple Method to Boost Performance of End-to-End Automatic Speech Recognition 1 1 2 Jul 2021 not harvested
HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units 11 2 14 Jun 2021 ran 0 of 9 samples (9 unverified)
Librispeech Transducer Model with Internal Language Model Prior Correction 2 2 7 Apr 2021 not harvested
SpeechStew: Simply Mix All Available Speech Recognition Data to Train One Large Neural Network 0 4 5 Apr 2021 not harvested
Transformer-based ASR Incorporating Time-reduction Layer and Fine-tuning with Self-Knowledge Distillation 0 1 17 Mar 2021 not harvested
Improving RNN Transducer Based ASR with Auxiliary Tasks 1 2 5 Nov 2020 not harvested
Self-training and Pre-training are Complementary for Speech Recognition 3 3 22 Oct 2020 not harvested
Pushing the Limits of Semi-Supervised Learning for Automatic Speech Recognition 1 2 20 Oct 2020 ran 2 of 2 samples (0 unverified; 2 pointer-only for licence)
wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations 25 3 20 Jun 2020 ran 2 of 9 samples (7 unverified; 2 pointer-only for licence)
ASAPP-ASR: Multistream CNN and Self-Attentive SRU for SOTA Speech Recognition 0 2 21 May 2020 not harvested
Iterative Pseudo-Labeling for Speech Recognition 1 2 19 May 2020 not harvested
Improved Noisy Student Training for Automatic Speech Recognition 1 2 19 May 2020 not harvested

The full list of 53 is in the JSON twin.

Dataset loaders archive 2025-07-28

4 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

CC BY 4.0

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • Kazakh Speech Corpus 2 (KSC2)
  • LibriSpeech test other
  • LibriSpeech test clean
  • CommonVoice (clean)
  • LibriSpeech and External
  • librispeech_asr
  • Librispeech (other)
  • Librispeech (clean)
  • LibriSpeech ASR
  • LibriSpeechLibri-Light test-othertest-other
  • LibriSpeech
  • LibriSpeech test-other
  • LibriSpeech test-clean

13 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections