Papers › Visual Speech Recognition for Multiple Languages in the Wild

Visual Speech Recognition for Multiple Languages in the Wild

26 Feb 2022arXiv:2202.13084archive 2025-07-28

Pingchuan Ma, Stavros Petridis, Maja Pantic

Visual speech recognition (VSR) aims to recognize the content of speech based on lip movements, without relying on the audio stream. Advances in deep learning and the availability of large audio-visual datasets have led to the development of much more accurate and robust VSR models than ever before. However, these advances are usually due to the larger training sets rather than the model design. Here we demonstrate that designing better models is equally as important as using larger training sets. We propose the addition of prediction-based auxiliary tasks to a VSR model, and highlight the importance of hyperparameter optimization and appropriate data augmentations. We show that such a model works for different languages and outperforms all previous methods trained on publicly available datasets by a large margin. It even outperforms models that were trained on non-publicly available datasets containing up to to 21 times more data. We show, furthermore, that using additional training data, even in other languages or with automatically generated transcriptions, results in further improvement.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

mpc001/Visual_Speech_Recognition_for_Multiple_Languages officialmentioned on GitHubpytorchNOASSERTION report
david-gimeno/lip-rtve mentioned on GitHubNOASSERTION report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Hyperparameter OptimizationLipreadingSpeech RecognitionVisual Speech Recognitionspeech-recognition

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Lipreading CMLR CTC/Attention CER 9.1% #1 of 5 Archive leaderboard report
Lipreading GRID corpus (mixed-speech) CTC/Attention Word Error Rate (WER) 1.2 #1 of 5 Archive leaderboard report
Lipreading LRS2 CTC/Attention (LRW+LRS2/3+AVSpeech) Word Error Rate (WER) 25.5 #7 of 25 Archive leaderboard report
Lipreading LRS2 CTC/Attention Word Error Rate (WER) 32.9 #15 of 25 Archive leaderboard report
Lipreading LRS3-TED CTC/Attention (LRW+LRS2/3+AVSpeech) Word Error Rate (WER) 31.5 #13 of 23 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections