Datasets › LRS2

LRS2 (Lip Reading Sentences 2)

Introduced by Joon Son Chung et al. in Lip Reading Sentences in the Wild1 Jan 2017 archive 2025-07-28

The Oxford-BBC Lip Reading Sentences 2 (LRS2) dataset is one of the largest publicly available datasets for lip reading sentences in-the-wild. The database consists of mainly news and talk shows from BBC programs. Each sentence is up to 100 characters in length. The training, validation and test sets are divided according to broadcast date. It is a challenging set since it contains thousands of speakers without speaker labels and large variation in head pose. The pre-training set contains 96,318 utterances, the training set contains 45,839 utterances, the validation set contains 1,082 utterances and the test set contains 1,242 utterances.

Source: Audio-visual Recognition of Overlapped speech for the LRS2 dataset Image Source: https://www.robots.ox.ac.uk/~vgg/data/lip_reading/lrs2.html

Benchmarks archive 2025-07-28

All 10 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

Papers archive 2025-07-28

26 shown of 26 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 115. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Unified Speech Recognition: A Single Model for Auditory, Visual, and Audiovisual Inputs 1 1 4 Nov 2024 ran 0 of 9 samples (9 unverified; 9 pointer-only for licence)
SyncVSR: Data-Efficient Visual Speech Recognition with End-to-End Crossmodal Audio Token Synchronization 1 3 18 Jun 2024 not harvested
Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation 1 2 14 Jun 2024 ran 5 of 18 samples (13 unverified; 18 pointer-only for licence)
TDFNet: An Efficient Audio-Visual Speech Separation Model with Top-down Fusion 1 3 25 Jan 2024 not harvested
ES3: Evolving Self-Supervised Learning of Robust Audio-Visual Speech Representations 0 6 1 Jan 2024 not harvested
Whispering LLaMA: A Cross-Modal Generative Error Correction Framework for Speech Recognition 1 1 10 Oct 2023 ran 4 of 5 samples (1 unverified)
RTFS-Net: Recurrent Time-Frequency Modelling for Efficient Audio-Visual Speech Separation 1 3 29 Sep 2023 ran 0 of 1 samples (1 unverified)
IIANet: An Intra- and Inter-Modality Attention Network for Audio-Visual Speech Separation 1 1 16 Aug 2023 ran 1 of 1 samples (0 unverified)
Auto-AVSR: Audio-Visual Speech Recognition with Automatic Labels 2 3 25 Mar 2023 ran 0 of 6 samples (6 unverified)
An Audio-Visual Speech Separation Model Inspired by Cortico-Thalamo-Cortical Circuits 2 1 21 Dec 2022 ran 2 of 3 samples (1 unverified)
Jointly Learning Visual and Auditory Speech Representations from Raw Data 1 2 12 Dec 2022 not harvested
MARLIN: Masked Autoencoder for facial video Representation LearnINg 1 1 12 Nov 2022 not harvested
Distinguishing Homophenes Using Multi-Head Visual-Audio Memory for Lip Reading 1 1 4 Apr 2022 not harvested
Visual Speech Recognition for Multiple Languages in the Wild 2 2 26 Feb 2022 not harvested
Leveraging Unimodal Self-Supervised Learning for Multimodal Audio-Visual Speech Recognition 1 3 24 Feb 2022 ran 0 of 6 samples (6 unverified)
Visual Keyword Spotting with Attention 1 1 29 Oct 2021 not harvested
Sub-word Level Lip Reading With Visual Attention 0 4 14 Oct 2021 not harvested
Image Shape Manipulation from a Single Augmented Training Sample 1 2 13 Sep 2021 not harvested
End-to-end Audio-visual Speech Recognition with Conformers 3 3 12 Feb 2021 not harvested
A Lip Sync Expert Is All You Need for Speech to Lip Generation In The Wild 4 2 23 Aug 2020 ran 1 of 2 samples (1 unverified; 2 pointer-only for licence)
Audio-visual Recognition of Overlapped speech for the LRS2 dataset 0 3 6 Jan 2020 not harvested
ASR is all you need: cross-modal distillation for lip reading 0 1 28 Nov 2019 not harvested
Hearing Lips: Improving Lip Reading by Distilling Speech Recognizers 1 1 26 Nov 2019 not harvested
Spatio-Temporal Fusion Based Convolutional Sequence Learning for Lip Reading 0 1 1 Oct 2019 not harvested
Audio-Visual Speech Recognition With A Hybrid CTC/Attention Architecture 0 3 28 Sep 2018 not harvested
Deep Audio-Visual Speech Recognition 4 6 6 Sep 2018 not harvested

Dataset loaders archive 2025-07-28

1 loader as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

Custom (non-commercial)

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • LRS2

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections