Papers › Audio-Visual Speech Recognition With A Hybrid CTC/Attention Architecture

Audio-Visual Speech Recognition With A Hybrid CTC/Attention Architecture

28 Sep 2018arXiv:1810.00108archive 2025-07-28

Stavros Petridis, Themos Stafylakis, Pingchuan Ma, Georgios Tzimiropoulos, Maja Pantic

Recent works in speech recognition rely either on connectionist temporal classification (CTC) or sequence-to-sequence models for character-level recognition. CTC assumes conditional independence of individual characters, whereas attention-based models can provide nonsequential alignments. Therefore, we could use a CTC loss in combination with an attention-based model in order to force monotonic alignments and at the same time get rid of the conditional independence assumption. In this paper, we use the recently proposed hybrid CTC/attention architecture for audio-visual recognition of speech in-the-wild. To the best of our knowledge, this is the first time that such a hybrid architecture architecture is used for audio-visual recognition of speech. We use the LRS2 database and show that the proposed audio-visual model leads to an 1.3% absolute decrease in word error rate over the audio-only model and achieves the new state-of-the-art performance on LRS2 database (7% word error rate). We also observe that the audio-visual model significantly outperforms the audio-based model (up to 32.9% absolute improvement in word error rate) for several different types of noise as the signal-to-noise ratio decreases.

PaperPDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Audio-Visual Speech RecognitionAutomatic Speech Recognition (ASR)LipreadingSpeech RecognitionVisual Speech Recognitionspeech-recognition

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Audio-Visual Speech Recognition LRS2 CTC/Attention Test WER 7.0 #6 of 8 Archive leaderboard report
Automatic Speech Recognition (ASR) LRS2 CTC/attention Test WER 8.2 #7 of 9 Archive leaderboard report
Lipreading LRS2 Hybrid CTC / Attention Word Error Rate (WER) 50 #21 of 25 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

CTC Loss

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections