Browse State-of-the-Art › Lip Reading
Lip Reading
49 papers with code · 3 benchmarks · 6 datasets archive 2025-07-28
Lip Reading is a task to infer the speech content in a video by using only the visual information, especially the lip movements. It has many crucial applications in practice, such as assisting audio-based speech recognition, biometric authentication and aiding hearing-impaired people.
Source: Mutual Information Maximization for Effective Lip Reading
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
3 leaderboard tables shown for this task, 3 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| GRID corpus (mixed-speech) (1 row) | Lip2Wav | Learning Individual Speaking Styles for Accurate Lip to Speech Synthesis | code | Syntology ran 1 of 9 samples · 8 unverified | Compare |
| LRW (1 row) | Lip2Wav | Learning Individual Speaking Styles for Accurate Lip to Speech Synthesis | code | Syntology ran 1 of 9 samples · 8 unverified | Compare |
| TCD-TIMIT corpus (mixed-speech) (1 row) | Lip2Wav | Learning Individual Speaking Styles for Accurate Lip to Speech Synthesis | code | Syntology ran 1 of 9 samples · 8 unverified | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
6 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
1 subtask in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 49 papers with code (153 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
6 Sep 2018 4 repositories listedThe goal of this work is to recognise phrases and sentences being spoken by a talking face, with or without the audio.
-
12 Mar 2017 4 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 1 pointer-only (licence)We propose an end-to-end deep learning architecture for word-level visual speech recognition.
-
12 Feb 2021 3 repositories listedIn this work, we present a hybrid CTC/Attention model based on a ResNet-18 and Convolution-augmented transformer (Conformer), that can be trained in an end-to-end manner.
-
9 Mar 2023 2 repositories listed Syntology ran 7 of 7 samples · 0 unverifiedHowever, despite researchers exploring cross-lingual translation techniques such as machine translation and audio speech translation to overcome language barriers, there is still a shortage of cross-lingual studies on…
-
5 Jan 2022 2 repositories listedThe lip-reading WER is further reduced to 26.
-
4 Dec 2020 2 repositories listedBiometric systems based on Machine learning and Deep learning are being extensively used as authentication mechanisms in resource-constrained environments like smartphones and other small computing devices.
-
23 Jan 2020 2 repositories listed Syntology ran 1 of 2 samples · 1 unverified · 2 pointer-only (licence)We present results on the largest publicly-available datasets for isolated word recognition in English and Mandarin, LRW and LRW1000, respectively.
-
16 Oct 2018 2 repositories listedIt has shown a large variation in this benchmark in several aspects, including the number of samples in each class, video resolution, lighting conditions, and speakers' attributes such as pose, age, gender, and make-up.
-
4 May 2025 1 repository listedConducted experiments confirm applicability of the proposed face ReID algorithm that is combining the concepts of face detection, face recognition and passive tracking-by-detection in order to achieve robust and…
-
11 Mar 2025 1 repository listedAutomatic CS Recognition (ACSR) refers to the AI-driven process of automatically recognizing hand gestures and lip movements in CS, converting them into text.
-
2 Sep 2024 1 repository listedTo address this challenge, speaker adaptive lip reading technologies have advanced by focusing on effectively adapting a lip reading model to target speakers in the visual modality.
-
9 May 2024 1 repository listedThe article introduces a novel audio-visual speech command recognition transformer (AVCRFormer) specifically designed for robust AVSR.
-
18 Apr 2024 1 repository listedThen we design a spatio-temporal fusion module based on temporal granularity alignment, where the global spatial features extracted from event frames, together with the local relative spatial and temporal features…
-
23 Feb 2024 1 repository listed Syntology ran 8 of 12 samples · 4 unverified · 12 pointer-only (licence)In this paper, we propose a novel framework, namely Visual Speech Processing incorporated with LLMs (VSP-LLM), to maximize the context modeling ability by bringing the overwhelming power of LLMs.
-
11 Dec 2023 1 repository listedOur method, which we call NEUral Text to ARticulate Talk (NEUTART), is a talking face generator that uses a joint audiovisual feature space, as well as speech-informed 3D facial reconstructions and a lip-reading loss…
-
23 Nov 2023 1 repository listedThe Lip Reading Sentences-3 (LRS3) benchmark has primarily been the focus of intense research in visual speech recognition (VSR) during the last few years.
-
8 Oct 2023 1 repository listedFor deep layers where both the speaker's features and the speech content features are all expressed well, we introduce the speaker-adaptive features to learn for suppressing the speech content irrelevant noise for…
-
19 Jun 2023 1 repository listed Syntology ran 3 of 3 samples · 0 unverified · 3 pointer-only (licence)To enhance the visual accuracy of generated lip movement while reducing the dependence on labeled data, we propose a novel framework SelfTalk, by involving self-supervision in a cross-modals network system to learn 3D…
-
10 Jun 2023 1 repository listedWe demonstrate that OpenSR enables modality transfer from one to any in three different settings (zero-, few- and full-shot), and achieves highly competitive zero-shot performance compared to the existing few-shot and…
-
5 Jun 2023 1 repository listedCued Speech (CS) is a multi-modal visual coding system combining lip reading with several hand cues at the phonetic level to make the spoken language visible to the hearing impaired.
-
5 Jun 2023 1 repository listed Syntology ran 5 of 8 samples · 3 unverifiedWe then condition a diffusion model on the video and use the extracted text through a classifier-guidance mechanism where a pre-trained ASR serves as the classifier.
-
29 Mar 2023 1 repository listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)To address the problem, we propose using a lip-reading expert to improve the intelligibility of the generated lip regions by penalizing the incorrect generation results.
-
31 Jan 2023 1 repository listed Syntology ran 3 of 4 samples · 1 unverifiedGenerating photo-realistic video portrait with arbitrary speech audio is a crucial problem in film-making and virtual reality.
-
16 Jan 2023 1 repository listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)Inspired by humans comprehending speech in a multi-modal manner, various audio-visual datasets have been constructed.
-
4 Jan 2023 1 repository listed Syntology ran 0 of 6 samples · 6 unverifiedWe improve previous lip reading methods using an Efficient Conformer back-end on top of a ResNet-18 visual front-end and by adding intermediate CTC losses between blocks.
-
7 Nov 2022 1 repository listedDeepfake technology has advanced a lot, but it is a double-sided sword for the community.
-
20 Sep 2022 1 repository listedThe powerful modeling capabilities of all-attention-based transformer architectures often cause overfitting and - for natural language processing tasks - lead to an implicitly learned internal language model in the…
-
3 Sep 2022 1 repository listedIn this paper, we systematically investigate the performance of state-of-the-art data augmentation approaches, temporal models and other training strategies, like self-distillation and using word boundary indicators.
-
9 Aug 2022 1 repository listedIn this paper, to remedy the performance degradation of lip reading model on unseen speakers, we propose a speaker-adaptive lip reading method, namely user-dependent padding.
-
4 Apr 2022 1 repository listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)By learning the interrelationship through the associative bridge, the proposed bridging framework is able to obtain the target modal representations inside the memory network, even with the source modal input only, and…
Syntology lines on 11 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections