Datasets › LRW

LRW (Lip Reading in the Wild)

Introduced in Lip Reading in the Wild1 Jan 2016 archive 2025-07-28

The Lip Reading in the Wild (LRW) dataset a large-scale audio-visual database that contains 500 different words from over 1,000 speakers. Each utterance has 29 frames, whose boundary is centered around the target word. The database is divided into training, validation and test sets. The training set contains at least 800 utterances for each class while the validation and test sets contain 50 utterances.

Source: Towards Pose-invariant Lip-Reading Image Source: https://www.robots.ox.ac.uk/~vgg/data/lip_reading/lrw1.html

Benchmarks archive 2025-07-28

All 8 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

Papers archive 2025-07-28

29 shown of 29 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 188. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
SyncVSR: Data-Efficient Visual Speech Recognition with End-to-End Crossmodal Audio Token Synchronization 1 4 18 Jun 2024 not harvested
Audio-Visual Speech Recognition based on Regulated Transformer and Spatio-Temporal Fusion Strategy for Driver Assistive Systems 1 2 9 May 2024 not harvested
Another Point of View on Visual Speech Recognition 0 1 20 Aug 2023 not harvested
Audio-Visual Speech and Gesture Recognition by Sensors of Mobile Devices 0 1 17 Feb 2023 not harvested
Training Strategies for Improved Lip-reading 1 1 3 Sep 2022 not harvested
Visual Speech Recognition in a Driver Assistance System 0 1 29 Aug 2022 not harvested
Accurate and Resource-Efficient Lipreading with Efficientnetv2 and Transformers 0 1 23 May 2022 not harvested
Multi-modality Associative Bridging through Memory: Speech Sound Recollected from Face Video 1 1 4 Apr 2022 ran 1 of 1 samples (0 unverified; 1 pointer-only for licence)
Distinguishing Homophenes Using Multi-Head Visual-Audio Memory for Lip Reading 1 1 4 Apr 2022 not harvested
Leveraging Unimodal Self-Supervised Learning for Multimodal Audio-Visual Speech Recognition 1 1 24 Feb 2022 ran 0 of 6 samples (6 unverified)
Visual Keyword Spotting with Attention 1 1 29 Oct 2021 not harvested
Adaptive Semantic-Spatio-Temporal Graph Convolutional Network for Lip Reading 0 1 16 Aug 2021 not harvested
Part-based Lipreading for Audio-Visual Speech Recognition 0 1 14 Dec 2020 not harvested
Learn an Effective Lip Reading Model without Pains 1 2 15 Nov 2020 not harvested
Lip Graph Assisted Audio-Visual Speech Recognition Using Bidirectional Synchronous Fusion 0 1 25 Oct 2020 not harvested
A Lip Sync Expert Is All You Need for Speech to Lip Generation In The Wild 4 2 23 Aug 2020 ran 1 of 2 samples (1 unverified; 2 pointer-only for licence)
Towards Practical Lipreading with Distilled and Efficient Models 1 1 13 Jul 2020 not harvested
SpotFast Networks with Memory Augmented Lateral Transformers for Lipreading 1 1 21 May 2020 not harvested
Learning Individual Speaking Styles for Accurate Lip to Speech Synthesis 1 2 17 May 2020 ran 1 of 9 samples (8 unverified)
Discriminative Multi-modality Speech Recognition 2 1 12 May 2020 not harvested
Mutual Information Maximization for Effective Lip Reading 1 1 13 Mar 2020 not harvested
Deformation Flow Based Two-Stream Network for Lip Reading 1 1 12 Mar 2020 not harvested
Pseudo-Convolutional Policy Gradient for Sequence-to-Sequence Lip-Reading 0 1 9 Mar 2020 not harvested
Can We Read Speech Beyond the Lips? Rethinking RoI Selection for Deep Visual Speech Recognition 1 1 6 Mar 2020 not harvested
Towards Automatic Face-to-Face Translation 1 1 1 Mar 2020 ran 2 of 2 samples (0 unverified)
Lipreading using Temporal Convolutional Networks 2 1 23 Jan 2020 ran 1 of 2 samples (1 unverified; 2 pointer-only for licence)
Multi-Grained Spatio-temporal Modeling for Lip-reading 0 1 30 Aug 2019 not harvested
End-to-end Audiovisual Speech Recognition 2 1 18 Feb 2018 not harvested
Combining Residual Networks with LSTMs for Lipreading 4 1 12 Mar 2017 ran 2 of 2 samples (0 unverified; 1 pointer-only for licence)

Dataset loaders archive 2025-07-28

1 loader as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

Custom (research-only, non-commercial, attribution)

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • LRW
  • Lip Reading in the Wild
  • Lipreading in the Wild

3 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections