Browse State-of-the-Art › Visual Speech Recognition
Visual Speech Recognition
62 papers with code · 2 benchmarks · 6 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
2 leaderboard tables shown for this task, 2 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| LRS3-TED (3 rows) | CTC/Attention | Auto-AVSR: Audio-Visual Speech Recognition with Automatic Labels | code | Syntology ran 0 of 6 samples · 6 unverified | Compare |
| LRS2 (2 rows) | VTP with more data | Sub-word Level Lip Reading With Visual Attention | — | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
6 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
1 subtask in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 62 papers with code (182 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
6 Sep 2018 4 repositories listedThe goal of this work is to recognise phrases and sentences being spoken by a talking face, with or without the audio.
-
12 Mar 2017 4 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 1 pointer-only (licence)We propose an end-to-end deep learning architecture for word-level visual speech recognition.
-
12 Feb 2021 3 repositories listedIn this work, we present a hybrid CTC/Attention model based on a ResNet-18 and Convolution-augmented transformer (Conformer), that can be trained in an end-to-end manner.
-
7 Jan 2024 2 repositories listedThis paper delineates the visual speech recognition (VSR) system introduced by the NPU-ASLP-LiAuto (Team 237) in the first Chinese Continuous Visual Speech Recognition Challenge (CNVSRC) 2023, engaging in the fixed and…
-
25 Mar 2023 2 repositories listed Syntology ran 0 of 6 samples · 6 unverifiedRecently, the performance of automatic, visual, and audio-visual speech recognition (ASR, VSR, and AV-ASR, respectively) has been substantially improved, mainly due to the use of larger models and training sets.
-
9 Mar 2023 2 repositories listed Syntology ran 7 of 7 samples · 0 unverifiedHowever, despite researchers exploring cross-lingual translation techniques such as machine translation and audio speech translation to overcome language barriers, there is still a shortage of cross-lingual studies on…
-
26 Feb 2022 2 repositories listedHowever, these advances are usually due to the larger training sets rather than the model design.
-
16 Oct 2018 2 repositories listedIt has shown a large variation in this benchmark in several aspects, including the number of samples in each class, video resolution, lighting conditions, and speakers' attributes such as pose, age, gender, and make-up.
-
20 May 2025 1 repository listedSign Language Recognition (SLR) systems primarily focus on manual gestures, but non-manual features such as mouth movements, specifically mouthing, provide valuable linguistic information.
-
6 May 2025 1 repository listedThey also enable strong performance in Visual Speech Recognition (VSR) with a WER of 20.
-
MMS-LLaMA: Efficient LLM-based Audio-Visual Speech Recognition with Minimal Multimodal Speech Tokens14 Mar 2025 1 repository listedAudio-Visual Speech Recognition (AVSR) achieves robust speech recognition in noisy environments by combining auditory and visual information.
-
8 Mar 2025 1 repository listedTo capture the wide spectrum of phonetic and linguistic diversity, we also introduce a Multilingual Audio-Visual Romanized Corpus (MARC) consisting of 2, 916 hours of audio-visual speech data across 82 languages, along…
-
9 Feb 2025 1 repository listedWe also introduce a multi-teacher ensemble method to distill the student, which receives audio-visual data as inputs.
-
3 Feb 2025 1 repository listedIn this work, we propose mWhisper-Flamingo for multilingual AVSR which combines the strengths of a pre-trained audio model (Whisper) and video model (AV-HuBERT).
-
1 Feb 2025 1 repository listedIn addition, a thorough ablation study is carried out, where it is studied how the different components that form the architecture influence the quality of speech recognition.
-
23 Jan 2025 1 repository listed Syntology ran 1 of 1 samples · 0 unverifiedSpecifically, we suggest a unimodal multi-task learning, which distills cross-modal knowledge and aligns the corrupted modalities, by predicting clean audio targets with corrupted videos, and clean video targets with…
-
3 Jan 2025 1 repository listedUnlike traditional Automatic Speech Recognition (ASR), Audio-Visual Speech Recognition (AVSR) takes audio and visual signals simultaneously to infer the transcription.
-
21 Oct 2024 1 repository listedThen, based on the temporal correspondence between audio and video, a frame-level local alignment loss is introduced to refine the global alignment, improving the utility of the audio information.
-
18 Sep 2024 1 repository listed Syntology ran 9 of 12 samples · 3 unverified · 12 pointer-only (licence)For example, in the audio and speech domains, an LLM can be equipped with (automatic) speech recognition (ASR) abilities by just concatenating the audio tokens, computed with an audio encoder, and the text tokens to…
-
5 Aug 2024 1 repository listedThis paper delineates the visual speech recognition (VSR) system introduced by the NPU-ASLP (Team 237) in the second Chinese Continuous Visual Speech Recognition Challenge (CNVSRC 2024), engaging in all four tracks,…
-
1 Aug 2024 1 repository listedIn this work, we present SynesLM, an unified model which can perform three multimodal language understanding tasks: audio-visual automatic speech recognition(AV-ASR) and visual-aided speech/machine translation(VST/VMT).
-
9 Jul 2024 1 repository listedRecent advances in Audio-Visual Speech Recognition (AVSR) have led to unprecedented achievements in the field, improving the robustness of this type of system in adverse, noisy environments.
-
4 Jul 2024 1 repository listed Syntology ran 10 of 13 samples · 3 unverifiedCross-modal attention modules are introduced to enrich video features with audio information so that speech variability can be taken into account when training on the video temporal dynamics.
-
18 Jun 2024 1 repository listedVisual Speech Recognition (VSR) stands at the intersection of computer vision and speech recognition, aiming to interpret spoken content from visual cues.
-
14 Jun 2024 1 repository listed Syntology ran 5 of 18 samples · 13 unverified · 18 pointer-only (licence)Since videos are harder to obtain than audio, the video training data of AVSR models is usually limited to a few thousand hours.
-
11 May 2024 1 repository listedIn this paper, we propose Watch Your Mouth, a novel method that leverages depth sensing to enable accurate silent speech recognition.
-
9 May 2024 1 repository listedThe article introduces a novel audio-visual speech command recognition transformer (AVCRFormer) specifically designed for robust AVSR.
-
7 Mar 2024 1 repository listedIn this paper, we investigate this contrasting phenomenon from the perspective of modality bias and reveal that an excessive modality bias on the audio caused by dropout is the underlying reason.
-
23 Feb 2024 1 repository listed Syntology ran 8 of 12 samples · 4 unverified · 12 pointer-only (licence)In this paper, we propose a novel framework, namely Visual Speech Processing incorporated with LLMs (VSP-LLM), to maximize the context modeling ability by bringing the overwhelming power of LLMs.
-
8 Feb 2024 1 repository listed Syntology ran 2 of 6 samples · 4 unverifiedRecent studies have successfully shown that large language models (LLMs) can be successfully used for generative error correction (GER) on top of the automatic speech recognition (ASR) output.
Syntology lines on 9 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections