Papers › End-to-end Audio-visual Speech Recognition with Conformers

End-to-end Audio-visual Speech Recognition with Conformers

12 Feb 2021arXiv:2102.06657archive 2025-07-28

Pingchuan Ma, Stavros Petridis, Maja Pantic

In this work, we present a hybrid CTC/Attention model based on a ResNet-18 and Convolution-augmented transformer (Conformer), that can be trained in an end-to-end manner. In particular, the audio and visual encoders learn to extract features directly from raw pixels and audio waveforms, respectively, which are then fed to conformers and then fusion takes place via a Multi-Layer Perceptron (MLP). The model learns to recognise characters using a combination of CTC and an attention mechanism. We show that end-to-end training, instead of using pre-computed visual features which is common in the literature, the use of a conformer, instead of a recurrent network, and the use of a transformer-based language model, significantly improve the performance of our model. We present results on the largest publicly available datasets for sentence-level speech recognition, Lip Reading Sentences 2 (LRS2) and Lip Reading Sentences 3 (LRS3), respectively. The results show that our proposed models raise the state-of-the-art performance by a large margin in audio-only, visual-only, and audio-visual experiments.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

mpc001/Visual_Speech_Recognition_for_Multiple_Languages mentioned on GitHubpytorchNOASSERTION report
mpc001/auto_avsr mentioned on GitHubpytorchApache-2.0 report
zziz/pwc pytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Audio-Visual Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage ModellingLip ReadingLipreadingSentenceSpeech RecognitionVisual Speech Recognitionspeech-recognition

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Audio-Visual Speech Recognition LRS2 End2end Conformer Test WER 3.7 #4 of 8 Archive leaderboard report
Audio-Visual Speech Recognition LRS3-TED Hyb-Conformer Word Error Rate (WER) 2.3 #9 of 12 Archive leaderboard report
Automatic Speech Recognition (ASR) LRS2 End2end Conformer Test WER 3.9 #4 of 9 Archive leaderboard report
Lipreading LRS2 Hybrid CTC / Attention Word Error Rate (WER) 39.1 #16 of 25 Archive leaderboard report
Lipreading LRS3-TED Hyb + Conformer Word Error Rate (WER) 43.3 #18 of 23 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections