{"url":"/dataset/librispeech","name":"LibriSpeech","full_name":null,"description_markdown":"The **LibriSpeech** corpus is a collection of approximately 1,000 hours of audiobooks that are a part of the LibriVox project. Most of the audiobooks come from the Project Gutenberg. The training data is split into 3 partitions of 100hr, 360hr, and 500hr sets while the dev and test data are split into the ’clean’ and ’other’ categories, respectively, depending upon how well or challenging Automatic Speech Recognition systems would perform against. Each of the dev and test sets is around 5hr in audio length. This corpus also provides the n-gram language models and the corresponding texts excerpted from the Project Gutenberg books, which contain 803M tokens and 977K unique words.\r\n\r\nSource: [State-of-the-art Speech Recognition using Multi-stream Self-attention with Dilated 1D Convolutions](https://arxiv.org/abs/1910.00716)","description_withheld":null,"homepage":"http://www.openslr.org/12","introduced_date":"2015-01-01","introduced_date_note":null,"introduced_by":{"paper":null,"title":"Librispeech: An ASR corpus based on public domain audio books","first_author":null,"url":"https://doi.org/10.1109/ICASSP.2015.7178964"},"license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"modalities":[{"name":"Audio","url":"/datasets/modality/audio"},{"name":"Speech","url":"/datasets/modality/speech"}],"tasks":[{"name":"Speech Recognition","url":"/task/speech-recognition","datasets_with_task":"/datasets/task/speech-recognition"},{"name":"Automatic Speech Recognition","url":"/task/automatic-speech-recognition-2","datasets_with_task":"/datasets/task/automatic-speech-recognition-2"},{"name":"Voice Conversion","url":"/task/voice-conversion","datasets_with_task":"/datasets/task/voice-conversion"},{"name":"Resynthesis","url":"/task/resynthesis","datasets_with_task":"/datasets/task/resynthesis"}],"languages":[{"name":"English","url":"/datasets/language/english"},{"name":"French","url":"/datasets/language/french"},{"name":"Spanish","url":"/datasets/language/spanish"},{"name":"Italian","url":"/datasets/language/italian"},{"name":"Japanese","url":"/datasets/language/japanese"},{"name":"Portuguese","url":"/datasets/language/portuguese"}],"variants":["Kazakh Speech Corpus 2 (KSC2)","LibriSpeech test other","LibriSpeech test clean","CommonVoice (clean)","LibriSpeech and External","librispeech_asr","Librispeech (other)","Librispeech (clean)","LibriSpeech ASR","LibriSpeechLibri-Light test-othertest-other","LibriSpeech","LibriSpeech test-other","LibriSpeech test-clean"],"data_loaders":[{"repo":"https://github.com/huggingface/datasets","url":"https://huggingface.co/datasets/librispeech_asr","frameworks":["tf","pytorch","jax"]},{"repo":"https://github.com/tensorflow/datasets","url":"https://www.tensorflow.org/datasets/catalog/librispeech","frameworks":["tf","jax"]},{"repo":"https://github.com/pytorch/audio","url":"https://pytorch.org/audio/stable/datasets.html#torchaudio.datasets.LIBRISPEECH","frameworks":["pytorch"]},{"repo":"https://gitlab.com/jaco-assistant/corcua","url":"https://gitlab.com/jaco-assistant/corcua","frameworks":[]}],"num_papers_in_archive":2361,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/speech-recognition-on-librispeech-test-clean","task":"Speech Recognition","dataset_variant":"LibriSpeech test-clean","rows":64,"metrics":["Word Error Rate (WER)"],"first_row_in_archive_order":{"model":"United Med ASR","paper":"/paper/high-precision-medical-speech-recognition","metrics":{"Word Error Rate (WER)":"0.985"},"code_links":[]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/speech-recognition-on-librispeech-test-other","task":"Speech Recognition","dataset_variant":"LibriSpeech test-other","rows":53,"metrics":["Word Error Rate (WER)"],"first_row_in_archive_order":{"model":"SAMBA ASR","paper":"/paper/samba-asr-state-of-the-art-speech-recognition","metrics":{"Word Error Rate (WER)":"2.48"},"code_links":[]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/voice-conversion-on-librispeech-test-clean","task":"Voice Conversion","dataset_variant":"LibriSpeech test-clean","rows":1,"metrics":["Character Error Rate (CER)","Equal Error Rate","Word Error Rate (WER)"],"first_row_in_archive_order":{"model":"kNN-VC (prematched HiFiGAN)","paper":"/paper/voice-conversion-with-just-nearest-neighbors","metrics":{"Character Error Rate (CER)":"2.96","Equal Error Rate":"37.15","Word Error Rate (WER)":"7.36"},"code_links":[{"title":"bshall/knn-vc","url":"https://github.com/bshall/knn-vc"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/speech-recognition-on-librispeech-other","task":"Speech Recognition","dataset_variant":"Librispeech (other)","rows":0,"metrics":["Test WER"],"first_row_in_archive_order":null,"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/let-ssms-be-convnets-state-space-modeling","title":"Let SSMs be ConvNets: State-space Modeling with Optimal Tensor Contractions","date":"2025-01-22","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/samba-asr-state-of-the-art-speech-recognition","title":"Samba-ASR: State-Of-The-Art Speech Recognition Leveraging Structured State-Space Models","date":"2025-01-06","rows_on_this_dataset":2,"code_links":0,"syntology":null},{"paper":"/paper/high-precision-medical-speech-recognition","title":"High-precision medical speech recognition through synthetic data and semantic correction: UNITED-MEDASR","date":"2024-11-24","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/cr-ctc-consistency-regularization-on-ctc-for","title":"CR-CTC: Consistency regularization on CTC for improved speech recognition","date":"2024-10-07","rows_on_this_dataset":4,"code_links":1,"syntology":null},{"paper":"/paper/fadam-adam-is-a-natural-gradient-optimizer","title":"FAdam: Adam is a natural gradient optimizer using diagonal empirical Fisher information","date":"2024-05-21","rows_on_this_dataset":2,"code_links":1,"syntology":null},{"paper":"/paper/graph-convolutions-enrich-the-self-attention","title":"Graph Convolutions Enrich the Self-Attention in Transformers!","date":"2023-12-07","rows_on_this_dataset":2,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":29,"samples_ran":19,"samples_unverified":10,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/qwen-audio-advancing-universal-audio","title":"Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models","date":"2023-11-14","rows_on_this_dataset":2,"code_links":2,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":7,"samples_ran":5,"samples_unverified":2,"pointer_only_for_licence":7,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/zipformer-a-faster-and-better-encoder-for","title":"Zipformer: A faster and better encoder for automatic speech recognition","date":"2023-10-17","rows_on_this_dataset":2,"code_links":1,"syntology":null},{"paper":"/paper/voice-conversion-with-just-nearest-neighbors","title":"Voice Conversion With Just Nearest Neighbors","date":"2023-05-30","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/multi-head-state-space-model-for-speech","title":"Multi-Head State Space Model for Speech Recognition","date":"2023-05-21","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/fast-conformer-with-linearly-scalable","title":"Fast Conformer with Linearly Scalable Attention for Efficient Speech Recognition","date":"2023-05-08","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/mt4ssl-boosting-self-supervised-speech","title":"MT4SSL: Boosting Self-Supervised Speech Representation Learning by Integrating Multiple Targets","date":"2022-11-14","rows_on_this_dataset":2,"code_links":1,"syntology":null},{"paper":"/paper/e-branchformer-branchformer-with-enhanced","title":"E-Branchformer: Branchformer with Enhanced merging for speech recognition","date":"2022-09-30","rows_on_this_dataset":2,"code_links":1,"syntology":null},{"paper":"/paper/squeezeformer-an-efficient-transformer-for","title":"Squeezeformer: An Efficient Transformer for Automatic Speech Recognition","date":"2022-06-02","rows_on_this_dataset":2,"code_links":4,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":49,"samples_ran":31,"samples_unverified":18,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/data2vec-a-general-framework-for-self-1","title":"data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language","date":"2022-02-07","rows_on_this_dataset":1,"code_links":12,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":6,"samples_ran":0,"samples_unverified":6,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/wavlm-large-scale-self-supervised-pre","title":"WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing","date":"2021-10-26","rows_on_this_dataset":2,"code_links":9,"syntology":null},{"paper":"/paper/w2v-bert-combining-contrastive-learning-and","title":"W2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre-Training","date":"2021-08-07","rows_on_this_dataset":2,"code_links":4,"syntology":null},{"paper":"/paper/amortized-neural-networks-for-low-latency","title":"Amortized Neural Networks for Low-Latency Speech Recognition","date":"2021-08-03","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/relaxed-attention-a-simple-method-to-boost","title":"Relaxed Attention: A Simple Method to Boost Performance of End-to-End Automatic Speech Recognition","date":"2021-07-02","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/hubert-self-supervised-speech-representation","title":"HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units","date":"2021-06-14","rows_on_this_dataset":2,"code_links":11,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":9,"samples_ran":0,"samples_unverified":9,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/librispeech-transducer-model-with-internal","title":"Librispeech Transducer Model with Internal Language Model Prior Correction","date":"2021-04-07","rows_on_this_dataset":2,"code_links":2,"syntology":null},{"paper":"/paper/speechstew-simply-mix-all-available-speech","title":"SpeechStew: Simply Mix All Available Speech Recognition Data to Train One Large Neural Network","date":"2021-04-05","rows_on_this_dataset":4,"code_links":0,"syntology":null},{"paper":"/paper/transformer-based-asr-incorporating-time","title":"Transformer-based ASR Incorporating Time-reduction Layer and Fine-tuning with Self-Knowledge Distillation","date":"2021-03-17","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/improving-rnn-transducer-based-asr-with","title":"Improving RNN Transducer Based ASR with Auxiliary Tasks","date":"2020-11-05","rows_on_this_dataset":2,"code_links":1,"syntology":null},{"paper":"/paper/self-training-and-pre-training-are","title":"Self-training and Pre-training are Complementary for Speech Recognition","date":"2020-10-22","rows_on_this_dataset":3,"code_links":3,"syntology":null},{"paper":"/paper/pushing-the-limits-of-semi-supervised","title":"Pushing the Limits of Semi-Supervised Learning for Automatic Speech Recognition","date":"2020-10-20","rows_on_this_dataset":2,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":2,"samples_ran":2,"samples_unverified":0,"pointer_only_for_licence":2,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/wav2vec-2-0-a-framework-for-self-supervised","title":"wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations","date":"2020-06-20","rows_on_this_dataset":3,"code_links":25,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":9,"samples_ran":2,"samples_unverified":7,"pointer_only_for_licence":2,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/asapp-asr-multistream-cnn-and-self-attentive","title":"ASAPP-ASR: Multistream CNN and Self-Attentive SRU for SOTA Speech Recognition","date":"2020-05-21","rows_on_this_dataset":2,"code_links":0,"syntology":null},{"paper":"/paper/iterative-pseudo-labeling-for-speech","title":"Iterative Pseudo-Labeling for Speech Recognition","date":"2020-05-19","rows_on_this_dataset":2,"code_links":1,"syntology":null},{"paper":"/paper/improved-noisy-student-training-for-automatic","title":"Improved Noisy Student Training for Automatic Speech Recognition","date":"2020-05-19","rows_on_this_dataset":2,"code_links":1,"syntology":null},{"paper":"/paper/fast-simpler-and-more-accurate-hybrid-asr","title":"Faster, Simpler and More Accurate Hybrid ASR Systems Using Wordpieces","date":"2020-05-19","rows_on_this_dataset":2,"code_links":0,"syntology":null},{"paper":"/paper/conformer-convolution-augmented-transformer","title":"Conformer: Convolution-augmented Transformer for Speech Recognition","date":"2020-05-16","rows_on_this_dataset":6,"code_links":25,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":7,"samples_ran":4,"samples_unverified":3,"pointer_only_for_licence":2,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/contextnet-improving-convolutional-neural","title":"ContextNet: Improving Convolutional Neural Networks for Automatic Speech Recognition with Global Context","date":"2020-05-07","rows_on_this_dataset":6,"code_links":6,"syntology":null},{"paper":"/paper/semi-supervised-speech-recognition-via-local","title":"Semi-Supervised Speech Recognition via Local Prior Matching","date":"2020-02-24","rows_on_this_dataset":3,"code_links":1,"syntology":null},{"paper":"/paper/end-to-end-asr-from-supervised-to-semi","title":"End-to-end ASR: from Supervised to Semi-Supervised Learning with Modern Architectures","date":"2019-11-19","rows_on_this_dataset":4,"code_links":1,"syntology":null},{"paper":"/paper/transformer-based-acoustic-modeling-for","title":"Transformer-based Acoustic Modeling for Hybrid Speech Recognition","date":"2019-10-22","rows_on_this_dataset":2,"code_links":0,"syntology":null},{"paper":"/paper/state-of-the-art-speech-recognition-using","title":"State-of-the-Art Speech Recognition Using Multi-Stream Self-Attention With Dilated 1D Convolutions","date":"2019-10-01","rows_on_this_dataset":2,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":1,"samples_ran":1,"samples_unverified":0,"pointer_only_for_licence":1,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/espresso-a-fast-end-to-end-neural-speech","title":"Espresso: A Fast End-to-end Neural Speech Recognition Toolkit","date":"2019-09-18","rows_on_this_dataset":2,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":1,"samples_ran":1,"samples_unverified":0,"pointer_only_for_licence":1,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/a-comparative-study-on-transformer-vs-rnn-in","title":"A Comparative Study on Transformer vs RNN in Speech Applications","date":"2019-09-13","rows_on_this_dataset":2,"code_links":2,"syntology":null},{"paper":"/paper/rwth-asr-systems-for-librispeech-hybrid-vs","title":"RWTH ASR Systems for LibriSpeech: Hybrid vs Attention -- w/o Data Augmentation","date":"2019-05-08","rows_on_this_dataset":2,"code_links":2,"syntology":null},{"paper":"/paper/specaugment-a-simple-data-augmentation-method","title":"SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition","date":"2019-04-18","rows_on_this_dataset":4,"code_links":30,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":18,"samples_ran":1,"samples_unverified":17,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/crf-based-single-stage-acoustic-modeling-with","title":"CRF-based Single-stage Acoustic Modeling with CTC Topology","date":"2019-04-16","rows_on_this_dataset":2,"code_links":1,"syntology":null},{"paper":"/paper/jasper-an-end-to-end-convolutional-neural","title":"Jasper: An End-to-End Convolutional Neural Acoustic Model","date":"2019-04-05","rows_on_this_dataset":4,"code_links":10,"syntology":null},{"paper":"/paper/model-unit-exploration-for-sequence-to","title":"On the Choice of Modeling Unit for Sequence-to-Sequence Speech Recognition","date":"2019-02-05","rows_on_this_dataset":1,"code_links":3,"syntology":null},{"paper":"/paper/fully-convolutional-speech-recognition","title":"Fully Convolutional Speech Recognition","date":"2018-12-17","rows_on_this_dataset":2,"code_links":0,"syntology":null},{"paper":"/paper/the-pytorch-kaldi-speech-recognition-toolkit","title":"The PyTorch-Kaldi Speech Recognition Toolkit","date":"2018-11-19","rows_on_this_dataset":1,"code_links":11,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":6,"samples_ran":5,"samples_unverified":1,"pointer_only_for_licence":6,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/snips-voice-platform-an-embedded-spoken","title":"Snips Voice Platform: an embedded Spoken Language Understanding system for private-by-design voice interfaces","date":"2018-05-25","rows_on_this_dataset":2,"code_links":16,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":15,"samples_ran":1,"samples_unverified":14,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/improved-training-of-end-to-end-attention","title":"Improved training of end-to-end attention models for speech recognition","date":"2018-05-08","rows_on_this_dataset":1,"code_links":14,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":12,"samples_ran":0,"samples_unverified":12,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/neural-network-language-modeling-with-letter","title":"Neural Network Language Modeling with Letter-based Features and Importance Sampling","date":"2018-04-15","rows_on_this_dataset":2,"code_links":0,"syntology":null},{"paper":"/paper/letter-based-speech-recognition-with-gated","title":"Letter-Based Speech Recognition with Gated ConvNets","date":"2017-12-22","rows_on_this_dataset":1,"code_links":2,"syntology":null},{"paper":"/paper/improving-end-to-end-speech-recognition-with-1","title":"Improving End-to-End Speech Recognition with Policy Learning","date":"2017-12-19","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/deep-speech-2-end-to-end-speech-recognition","title":"Deep Speech 2: End-to-End Speech Recognition in English and Mandarin","date":"2015-12-08","rows_on_this_dataset":2,"code_links":35,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":39,"samples_ran":2,"samples_unverified":37,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/quartznet-deep-automatic-speech-recognition","title":"QuartzNet: Deep Automatic Speech Recognition with 1D Time-Channel Separable Convolutions","date":null,"rows_on_this_dataset":2,"code_links":15,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":12,"samples_ran":3,"samples_unverified":9,"pointer_only_for_licence":2,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":16,"samples_harvested":222,"samples_ran":77,"samples_unverified":145,"pointer_only_for_licence":23,"papers_with_no_sample_that_ran":3,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}