Browse State-of-the-Art › Audio-Visual Speech Recognition
Audio-Visual Speech Recognition
42 papers with code · 4 benchmarks · 7 datasets archive 2025-07-28
Audio-visual speech recognition is the task of transcribing a paired audio and visual stream into text.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
4 leaderboard tables shown for this task, 4 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| LRS3-TED (12 rows) | MMS-LLaMA | MMS-LLaMA: Efficient LLM-based Audio-Visual Speech Recognition... | code | — | Compare |
| LRS2 (8 rows) | Whisper-Flamingo | Whisper-Flamingo: Integrating Visual Features into Whisper for... | code | Syntology ran 5 of 18 samples · 13 unverified | Compare |
| LRW (3 rows) | AVCRFormer | Audio-Visual Speech Recognition based on Regulated Transformer and... | code | — | Compare |
| CAS-VSR-S101 (1 row) | ES³ Base* | ES3: Evolving Self-Supervised Learning of Robust Audio-Visual... | — | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
7 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
1 subtask in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 42 papers with code (100 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
6 Sep 2018 4 repositories listedThe goal of this work is to recognise phrases and sentences being spoken by a talking face, with or without the audio.
-
12 Feb 2021 3 repositories listedIn this work, we present a hybrid CTC/Attention model based on a ResNet-18 and Convolution-augmented transformer (Conformer), that can be trained in an end-to-end manner.
-
25 Mar 2023 2 repositories listed Syntology ran 0 of 6 samples · 6 unverifiedRecently, the performance of automatic, visual, and audio-visual speech recognition (ASR, VSR, and AV-ASR, respectively) has been substantially improved, mainly due to the use of larger models and training sets.
-
12 May 2020 2 repositories listedVision is often used as a complementary modality for audio speech recognition (ASR), especially in the noisy environment where performance of solo audio modality significantly deteriorates.
-
6 May 2025 1 repository listedThey also enable strong performance in Visual Speech Recognition (VSR) with a WER of 20.
-
MMS-LLaMA: Efficient LLM-based Audio-Visual Speech Recognition with Minimal Multimodal Speech Tokens14 Mar 2025 1 repository listedAudio-Visual Speech Recognition (AVSR) achieves robust speech recognition in noisy environments by combining auditory and visual information.
-
8 Mar 2025 1 repository listedTo capture the wide spectrum of phonetic and linguistic diversity, we also introduce a Multilingual Audio-Visual Romanized Corpus (MARC) consisting of 2, 916 hours of audio-visual speech data across 82 languages, along…
-
9 Feb 2025 1 repository listedWe also introduce a multi-teacher ensemble method to distill the student, which receives audio-visual data as inputs.
-
3 Feb 2025 1 repository listedIn this work, we propose mWhisper-Flamingo for multilingual AVSR which combines the strengths of a pre-trained audio model (Whisper) and video model (AV-HuBERT).
-
23 Jan 2025 1 repository listed Syntology ran 1 of 1 samples · 0 unverifiedSpecifically, we suggest a unimodal multi-task learning, which distills cross-modal knowledge and aligns the corrupted modalities, by predicting clean audio targets with corrupted videos, and clean video targets with…
-
3 Jan 2025 1 repository listedUnlike traditional Automatic Speech Recognition (ASR), Audio-Visual Speech Recognition (AVSR) takes audio and visual signals simultaneously to infer the transcription.
-
18 Sep 2024 1 repository listed Syntology ran 9 of 12 samples · 3 unverified · 12 pointer-only (licence)For example, in the audio and speech domains, an LLM can be equipped with (automatic) speech recognition (ASR) abilities by just concatenating the audio tokens, computed with an audio encoder, and the text tokens to…
-
1 Aug 2024 1 repository listedIn this work, we present SynesLM, an unified model which can perform three multimodal language understanding tasks: audio-visual automatic speech recognition(AV-ASR) and visual-aided speech/machine translation(VST/VMT).
-
9 Jul 2024 1 repository listedRecent advances in Audio-Visual Speech Recognition (AVSR) have led to unprecedented achievements in the field, improving the robustness of this type of system in adverse, noisy environments.
-
4 Jul 2024 1 repository listed Syntology ran 10 of 13 samples · 3 unverifiedCross-modal attention modules are introduced to enrich video features with audio information so that speech variability can be taken into account when training on the video temporal dynamics.
-
14 Jun 2024 1 repository listed Syntology ran 5 of 18 samples · 13 unverified · 18 pointer-only (licence)Since videos are harder to obtain than audio, the video training data of AVSR models is usually limited to a few thousand hours.
-
9 May 2024 1 repository listedThe article introduces a novel audio-visual speech command recognition transformer (AVCRFormer) specifically designed for robust AVSR.
-
7 Mar 2024 1 repository listedIn this paper, we investigate this contrasting phenomenon from the perspective of modality bias and reveal that an excessive modality bias on the audio caused by dropout is the underlying reason.
-
8 Feb 2024 1 repository listed Syntology ran 2 of 6 samples · 4 unverifiedRecent studies have successfully shown that large language models (LLMs) can be successfully used for generative error correction (GER) on top of the automatic speech recognition (ASR) output.
-
7 Jan 2024 1 repository listed Syntology ran 6 of 9 samples · 3 unverified · 9 pointer-only (licence)Considering that visual information helps to improve speech recognition performance in noisy scenes, in this work we propose a multichannel multi-modal speech self-supervised learning framework AV-wav2vec2, which…
-
29 Sep 2023 1 repository listed Syntology ran 0 of 1 samples · 1 unverifiedThis is the first time-frequency domain audio-visual speech separation method to outperform all contemporary time-domain counterparts.
-
14 Aug 2023 1 repository listedIn this paper, we propose two novel techniques to improve audio-visual speech recognition (AVSR) under a pre-training and fine-tuning training framework.
-
18 Jun 2023 1 repository listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)In this paper, we aim to learn the shared representations across modalities to bridge their gap.
-
18 Jun 2023 1 repository listed Syntology ran 1 of 2 samples · 1 unverified · 2 pointer-only (licence)In this work, we investigate the noise-invariant visual modality to strengthen robustness of AVSR, which can adapt to any testing noises while without dependence on noisy training data, a.
-
10 Jun 2023 1 repository listedWe demonstrate that OpenSR enables modality transfer from one to any in three different settings (zero-, few- and full-shot), and achieves highly competitive zero-shot performance compared to the existing few-shot and…
-
4 Jun 2023 1 repository listedAudio-visual speech recognition (AVSR) gains increasing attention from researchers as an important part of human-computer interaction.
-
18 May 2023 1 repository listedWe investigate the emergent abilities of the recently proposed web-scale speech model Whisper, by adapting it to unseen tasks with prompt engineering.
-
16 May 2023 1 repository listedHowever, most existing AVSR approaches simply fuse the audio and visual features by concatenation, without explicit interactions to capture the deep correlations between them, which results in sub-optimal multimodal…
-
15 Mar 2023 1 repository listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)Thus, we firstly analyze that the previous AVSR models are not indeed robust to the corruption of multimodal input streams, the audio and the visual inputs, compared to uni-modal models.
-
1 Mar 2023 1 repository listedWe introduce MuAViC, a multilingual audio-visual corpus for robust speech recognition and robust speech-to-text translation providing 1200 hours of audio-visual speech in 9 languages.
Syntology lines on 11 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections