Datasets › LRS3-TED

LRS3-TED

Introduced by Triantafyllos Afouras et al. in LRS3-TED: a large-scale dataset for visual speech recognition archive 2025-07-28

LRS3-TED is a multi-modal dataset for visual and audio-visual speech recognition. It includes face tracks from over 400 hours of TED and TEDx videos, along with the corresponding subtitles and word alignment boundaries. The new dataset is substantially larger in scale compared to other public datasets that are available for general research.

Source: LRS3-TED: a large-scale dataset for visual speech recognition

Benchmarks archive 2025-07-28

All 7 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

Papers archive 2025-07-28

26 shown of 26 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 63. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
MMS-LLaMA: Efficient LLM-based Audio-Visual Speech Recognition with Minimal Multimodal Speech Tokens 1 1 14 Mar 2025 not harvested
Zero-AVSR: Zero-Shot Audio-Visual Speech Recognition with LLMs by Learning Language-Agnostic Speech Representations 1 1 8 Mar 2025 not harvested
Audio-Visual Representation Learning via Knowledge Distillation from Speech Foundation Models 1 3 9 Feb 2025 not harvested
Unified Speech Recognition: A Single Model for Auditory, Visual, and Audiovisual Inputs 1 2 4 Nov 2024 ran 0 of 9 samples (9 unverified; 9 pointer-only for licence)
Large Language Models are Strong Audio-Visual Speech Recognition Learners 1 2 18 Sep 2024 ran 9 of 12 samples (3 unverified; 12 pointer-only for licence)
SyncVSR: Data-Efficient Visual Speech Recognition with End-to-End Crossmodal Audio Token Synchronization 1 2 18 Jun 2024 not harvested
Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation 1 2 14 Jun 2024 ran 5 of 18 samples (13 unverified; 18 pointer-only for licence)
Where Visual Speech Meets Language: VSP-LLM Framework for Efficient and Context-Aware Visual Speech Processing 1 1 23 Feb 2024 ran 8 of 12 samples (4 unverified; 12 pointer-only for licence)
ES3: Evolving Self-Supervised Learning of Robust Audio-Visual Speech Representations 0 2 1 Jan 2024 not harvested
GestSync: Determining who is speaking without a talking head 1 1 8 Oct 2023 not harvested
Auto-AVSR: Audio-Visual Speech Recognition with Automatic Labels 2 4 25 Mar 2023 ran 0 of 6 samples (6 unverified)
Conformers are All You Need for Visual Speech Recognition 0 1 17 Feb 2023 not harvested
Jointly Learning Visual and Auditory Speech Representations from Raw Data 1 3 12 Dec 2022 not harvested
Relaxed Attention for Transformer Models 1 1 20 Sep 2022 not harvested
Visual Speech Recognition for Multiple Languages in the Wild 2 1 26 Feb 2022 not harvested
Robust Self-Supervised Audio-Visual Speech Recognition 1 1 5 Jan 2022 ran 1 of 1 samples (0 unverified; 1 pointer-only for licence)
Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction 2 2 5 Jan 2022 not harvested
Visual Keyword Spotting with Attention 1 1 29 Oct 2021 not harvested
Sub-word Level Lip Reading With Visual Attention 0 4 14 Oct 2021 not harvested
End-to-end Audio-visual Speech Recognition with Conformers 3 2 12 Feb 2021 not harvested
Discriminative Multi-modality Speech Recognition 2 2 12 May 2020 not harvested
ASR is all you need: cross-modal distillation for lip reading 0 1 28 Nov 2019 not harvested
Recurrent Neural Network Transducer for Audio-Visual Speech Recognition 1 2 8 Nov 2019 not harvested
Spatio-Temporal Fusion Based Convolutional Sequence Learning for Lip Reading 0 1 1 Oct 2019 not harvested
Deep Audio-Visual Speech Recognition 4 2 6 Sep 2018 not harvested
Large-Scale Visual Speech Recognition 0 1 13 Jul 2018 not harvested

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

Creative Commons BY-NC-ND 4.0 license

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • LRS3-TED

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections