{"url":"/dataset/lrs2","name":"LRS2","full_name":"Lip Reading Sentences 2","description_markdown":"The Oxford-BBC **Lip Reading Sentences 2** (**LRS2**) dataset is one of the largest publicly available datasets for lip reading sentences in-the-wild. The database consists of mainly news and talk shows from BBC programs. Each sentence is up to 100 characters in length. The training, validation and test sets are divided according to broadcast date. It is a challenging set since it contains thousands of speakers without speaker labels and large variation in head pose. The pre-training set contains 96,318 utterances, the training set contains 45,839 utterances, the validation set contains 1,082 utterances and the test set contains 1,242 utterances.\r\n\r\nSource: [Audio-visual Recognition of Overlapped speech for the LRS2 dataset](https://arxiv.org/abs/2001.01656)\r\nImage Source: [https://www.robots.ox.ac.uk/~vgg/data/lip_reading/lrs2.html](https://www.robots.ox.ac.uk/~vgg/data/lip_reading/lrs2.html)","description_withheld":null,"homepage":"https://www.robots.ox.ac.uk/~vgg/data/lip_reading/lrs2.html","introduced_date":"2017-01-01","introduced_date_note":null,"introduced_by":{"paper":"/paper/lip-reading-sentences-in-the-wild","title":"Lip Reading Sentences in the Wild","first_author":"Joon Son Chung","url":null},"license":{"name":"Custom (non-commercial)","url":"https://www.robots.ox.ac.uk/~vgg/data/lip_reading/lrs2.html"},"modalities":[{"name":"Videos","url":"/datasets/modality/videos"},{"name":"Texts","url":"/datasets/modality/texts"},{"name":"Audio","url":"/datasets/modality/audio"}],"tasks":[{"name":"Speech Recognition","url":"/task/speech-recognition","datasets_with_task":"/datasets/task/speech-recognition"},{"name":"Speech Separation","url":"/task/speech-separation","datasets_with_task":"/datasets/task/speech-separation"},{"name":"Automatic Speech Recognition (ASR)","url":"/task/automatic-speech-recognition","datasets_with_task":"/datasets/task/automatic-speech-recognition"},{"name":"Visual Speech Recognition","url":"/task/visual-speech-recognition","datasets_with_task":"/datasets/task/visual-speech-recognition"},{"name":"Lipreading","url":"/task/lipreading","datasets_with_task":"/datasets/task/lipreading"},{"name":"Audio-Visual Speech Recognition","url":"/task/audio-visual-speech-recognition","datasets_with_task":"/datasets/task/audio-visual-speech-recognition"},{"name":"Unconstrained Lip-synchronization","url":"/task/lip-sync","datasets_with_task":"/datasets/task/lip-sync"},{"name":"Visual Keyword Spotting","url":"/task/visual-keyword-spotting","datasets_with_task":"/datasets/task/visual-keyword-spotting"},{"name":"Landmark-based Lipreading","url":"/task/landmark-based-lipreading","datasets_with_task":"/datasets/task/landmark-based-lipreading"},{"name":"Image Manipulation","url":"/task/image-manipulation","datasets_with_task":"/datasets/task/image-manipulation"}],"languages":[],"variants":["LRS2"],"data_loaders":[{"repo":"https://github.com/Rudrabha/Wav2Lip","url":"https://github.com/Rudrabha/Wav2Lip?tab=readme-ov-file","frameworks":["pytorch"]}],"num_papers_in_archive":115,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/lipreading-on-lrs2","task":"Lipreading","dataset_variant":"LRS2","rows":25,"metrics":["Word Error Rate (WER)"],"first_row_in_archive_order":{"model":"Auto-AVSR","paper":"/paper/auto-avsr-audio-visual-speech-recognition","metrics":{"Word Error Rate (WER)":"14.6"},"code_links":[{"title":"mpc001/auto_avsr","url":"https://github.com/mpc001/auto_avsr"},{"title":"umbertocappellazzo/llama-avsr","url":"https://github.com/umbertocappellazzo/llama-avsr"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/automatic-speech-recognition-on-lrs2","task":"Automatic Speech Recognition (ASR)","dataset_variant":"LRS2","rows":9,"metrics":["Test WER"],"first_row_in_archive_order":{"model":"Whisper","paper":"/paper/whisper-flamingo-integrating-visual-features","metrics":{"Test WER":"1.3"},"code_links":[{"title":"roudimit/whisper-flamingo","url":"https://github.com/roudimit/whisper-flamingo"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/audio-visual-speech-recognition-on-lrs2","task":"Audio-Visual Speech Recognition","dataset_variant":"LRS2","rows":8,"metrics":["Test WER"],"first_row_in_archive_order":{"model":"Whisper-Flamingo","paper":"/paper/whisper-flamingo-integrating-visual-features","metrics":{"Test WER":"1.4"},"code_links":[{"title":"roudimit/whisper-flamingo","url":"https://github.com/roudimit/whisper-flamingo"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/speech-separation-on-lrs2","task":"Speech Separation","dataset_variant":"LRS2","rows":8,"metrics":["SI-SNRi","SDRi","PESQ","STOI"],"first_row_in_archive_order":{"model":"IIANet","paper":"/paper/scanet-a-self-and-cross-attention-network-for","metrics":{"SDRi":"16.6","SI-SNRi":"16.4"},"code_links":[{"title":"JusperLee/IIANet","url":"https://github.com/JusperLee/IIANet"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/lip-sync-on-lrs2","task":"Unconstrained Lip-synchronization","dataset_variant":"LRS2","rows":3,"metrics":["FID","LSE-D","LSE-C"],"first_row_in_archive_order":{"model":"Wav2Lip + ViT + MARLIN","paper":"/paper/marlin-masked-autoencoder-for-facial-video","metrics":{"FID":"3.452","LSE-C":"5.528","LSE-D":"7.127"},"code_links":[{"title":"ControlNet/MARLIN","url":"https://github.com/ControlNet/MARLIN"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/image-manipulation-on-lrs2","task":"Image Manipulation","dataset_variant":"LRS2","rows":2,"metrics":["LPIPS (S1)","LPIPS (S2)","LPIPS (S3)","LPIPS (S4)","LPIPS (S5)","SIFID (S1)","SIFID (S2)","SIFID (S3)","SIFID (S4)","SIFID (S5)"],"first_row_in_archive_order":{"model":"TPS","paper":"/paper/image-shape-manipulation-from-a-single","metrics":{"LPIPS (S1)":"0.12","LPIPS (S2)":"0.21","LPIPS (S3)":"0.1","LPIPS (S4)":"0.22","LPIPS (S5)":"0.14","SIFID (S1)":"0.07","SIFID (S2)":"0.12","SIFID (S3)":"0.04","SIFID (S4)":"0.12","SIFID (S5)":"0.06"},"code_links":[{"title":"eliahuhorwitz/DeepSIM","url":"https://github.com/eliahuhorwitz/DeepSIM"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/visual-speech-recognition-on-lrs2","task":"Visual Speech Recognition","dataset_variant":"LRS2","rows":2,"metrics":["Word Error Rate (WER)"],"first_row_in_archive_order":{"model":"VTP with more data","paper":"/paper/sub-word-level-lip-reading-with-visual","metrics":{"Word Error Rate (WER)":"22.6"},"code_links":[]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/landmark-based-lipreading-on-lrs2","task":"Landmark-based Lipreading","dataset_variant":"LRS2","rows":1,"metrics":["Word Error Rate (WER)"],"first_row_in_archive_order":{"model":"SyncVSR","paper":"/paper/syncvsr-data-efficient-visual-speech","metrics":{"Word Error Rate (WER)":"74.6"},"code_links":[{"title":"KAIST-AILab/SyncVSR","url":"https://github.com/KAIST-AILab/SyncVSR"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/speech-recognition-on-lrs2","task":"Speech Recognition","dataset_variant":"LRS2","rows":1,"metrics":["Word Error Rate (WER)"],"first_row_in_archive_order":{"model":"RAVEn Large","paper":"/paper/jointly-learning-visual-and-auditory-speech","metrics":{"Word Error Rate (WER)":"2.1"},"code_links":[{"title":"ahaliassos/raven","url":"https://github.com/ahaliassos/raven"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/visual-keyword-spotting-on-lrs2","task":"Visual Keyword Spotting","dataset_variant":"LRS2","rows":1,"metrics":["Top-1 Accuracy","Top-5 Accuracy","mAP","mAP IOU@0.5"],"first_row_in_archive_order":{"model":"Transpotter","paper":"/paper/visual-keyword-spotting-with-attention","metrics":{"Top-1 Accuracy":"65","Top-5 Accuracy":"87.1","mAP":"69.2","mAP IOU@0.5":"68.3"},"code_links":[{"title":"prajwalkr/transpotter","url":"https://github.com/prajwalkr/transpotter"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/unified-speech-recognition-a-single-model-for","title":"Unified Speech Recognition: A Single Model for Auditory, Visual, and Audiovisual Inputs","date":"2024-11-04","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":9,"samples_ran":0,"samples_unverified":9,"pointer_only_for_licence":9,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/syncvsr-data-efficient-visual-speech","title":"SyncVSR: Data-Efficient Visual Speech Recognition with End-to-End Crossmodal Audio Token Synchronization","date":"2024-06-18","rows_on_this_dataset":3,"code_links":1,"syntology":null},{"paper":"/paper/whisper-flamingo-integrating-visual-features","title":"Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation","date":"2024-06-14","rows_on_this_dataset":2,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":18,"samples_ran":5,"samples_unverified":13,"pointer_only_for_licence":18,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/tdfnet-an-efficient-audio-visual-speech","title":"TDFNet: An Efficient Audio-Visual Speech Separation Model with Top-down Fusion","date":"2024-01-25","rows_on_this_dataset":3,"code_links":1,"syntology":null},{"paper":"/paper/es3-evolving-self-supervised-learning-of","title":"ES3: Evolving Self-Supervised Learning of Robust Audio-Visual Speech Representations","date":"2024-01-01","rows_on_this_dataset":6,"code_links":0,"syntology":null},{"paper":"/paper/whispering-llama-a-cross-modal-generative","title":"Whispering LLaMA: A Cross-Modal Generative Error Correction Framework for Speech Recognition","date":"2023-10-10","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":5,"samples_ran":4,"samples_unverified":1,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/rtfs-net-recurrent-time-frequency-modelling","title":"RTFS-Net: Recurrent Time-Frequency Modelling for Efficient Audio-Visual Speech Separation","date":"2023-09-29","rows_on_this_dataset":3,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":1,"samples_ran":0,"samples_unverified":1,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/scanet-a-self-and-cross-attention-network-for","title":"IIANet: An Intra- and Inter-Modality Attention Network for Audio-Visual Speech Separation","date":"2023-08-16","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":1,"samples_ran":1,"samples_unverified":0,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/auto-avsr-audio-visual-speech-recognition","title":"Auto-AVSR: Audio-Visual Speech Recognition with Automatic Labels","date":"2023-03-25","rows_on_this_dataset":3,"code_links":2,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":6,"samples_ran":0,"samples_unverified":6,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/an-audio-visual-speech-separation-model","title":"An Audio-Visual Speech Separation Model Inspired by Cortico-Thalamo-Cortical Circuits","date":"2022-12-21","rows_on_this_dataset":1,"code_links":2,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":3,"samples_ran":2,"samples_unverified":1,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/jointly-learning-visual-and-auditory-speech","title":"Jointly Learning Visual and Auditory Speech Representations from Raw Data","date":"2022-12-12","rows_on_this_dataset":2,"code_links":1,"syntology":null},{"paper":"/paper/marlin-masked-autoencoder-for-facial-video","title":"MARLIN: Masked Autoencoder for facial video Representation LearnINg","date":"2022-11-12","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/distinguishing-homophenes-using-multi-head-1","title":"Distinguishing Homophenes Using Multi-Head Visual-Audio Memory for Lip Reading","date":"2022-04-04","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/visual-speech-recognition-for-multiple","title":"Visual Speech Recognition for Multiple Languages in the Wild","date":"2022-02-26","rows_on_this_dataset":2,"code_links":2,"syntology":null},{"paper":"/paper/leveraging-uni-modal-self-supervised-learning-1","title":"Leveraging Unimodal Self-Supervised Learning for Multimodal Audio-Visual Speech Recognition","date":"2022-02-24","rows_on_this_dataset":3,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":6,"samples_ran":0,"samples_unverified":6,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/visual-keyword-spotting-with-attention","title":"Visual Keyword Spotting with Attention","date":"2021-10-29","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/sub-word-level-lip-reading-with-visual","title":"Sub-word Level Lip Reading With Visual Attention","date":"2021-10-14","rows_on_this_dataset":4,"code_links":0,"syntology":null},{"paper":"/paper/image-shape-manipulation-from-a-single","title":"Image Shape Manipulation from a Single Augmented Training Sample","date":"2021-09-13","rows_on_this_dataset":2,"code_links":1,"syntology":null},{"paper":"/paper/end-to-end-audio-visual-speech-recognition","title":"End-to-end Audio-visual Speech Recognition with Conformers","date":"2021-02-12","rows_on_this_dataset":3,"code_links":3,"syntology":null},{"paper":"/paper/a-lip-sync-expert-is-all-you-need-for-speech","title":"A Lip Sync Expert Is All You Need for Speech to Lip Generation In The Wild","date":"2020-08-23","rows_on_this_dataset":2,"code_links":4,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":2,"samples_ran":1,"samples_unverified":1,"pointer_only_for_licence":2,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/audio-visual-recognition-of-overlapped-speech","title":"Audio-visual Recognition of Overlapped speech for the LRS2 dataset","date":"2020-01-06","rows_on_this_dataset":3,"code_links":0,"syntology":null},{"paper":"/paper/asr-is-all-you-need-cross-modal-distillation","title":"ASR is all you need: cross-modal distillation for lip reading","date":"2019-11-28","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/hearing-lips-improving-lip-reading-by","title":"Hearing Lips: Improving Lip Reading by Distilling Speech Recognizers","date":"2019-11-26","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/spatio-temporal-fusion-based-convolutional","title":"Spatio-Temporal Fusion Based Convolutional Sequence Learning for Lip Reading","date":"2019-10-01","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/audio-visual-speech-recognition-with-a-hybrid","title":"Audio-Visual Speech Recognition With A Hybrid CTC/Attention Architecture","date":"2018-09-28","rows_on_this_dataset":3,"code_links":0,"syntology":null},{"paper":"/paper/deep-audio-visual-speech-recognition","title":"Deep Audio-Visual Speech Recognition","date":"2018-09-06","rows_on_this_dataset":6,"code_links":4,"syntology":null}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":9,"samples_harvested":51,"samples_ran":13,"samples_unverified":38,"pointer_only_for_licence":29,"papers_with_no_sample_that_ran":4,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}