Browse State-of-the-Art › Visual Speech Recognition › Papers, page 2
Visual Speech Recognition
Papers archive 2025-07-28
archive papers tagged: 182 · with a code link: 62 · where Syntology ran a sample: 19 (16 with a run with no instrument failure, 3 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (19 of 182 tagged: 16 with a run with no instrument failure, 3 where every run was a failure of Syntology's instrument)
Page 2 of 2: papers 101 to 182 of 182, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
11 Aug 2023 0 repositories listed Syntology 17 ran (of which 0 constructed an object rather than computing a result; 17 with no instrument failure: 1 honoured, 0 violated, 16 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 19 harvested samples) · 19 pointer-only (licence)
-
SparseVSR: Lightweight and Noise Robust Visual Speech Recognition10 Jul 2023 0 repositories listed
-
Automated Speaker Independent Visual Speech Recognition: A Comprehensive Survey14 Jun 2023 0 repositories listed
-
Improving the Gap in Visual Speech Recognition Between Normal and Silent Speech Based on Metric Learning23 May 2023 0 repositories listed
-
Multi-Temporal Lip-Audio Memory for Visual Speech Recognition8 May 2023 0 repositories listed
-
Deep Learning-based Spatio Temporal Facial Feature Visual Speech Recognition30 Apr 2023 0 repositories listed
-
SynthVSR: Scaling Up Visual Speech Recognition With Synthetic Supervision30 Mar 2023 0 repositories listed
-
The NPU-ASLP System for Audio-Visual Speech Recognition in MISP 2022 Challenge11 Mar 2023 0 repositories listed
-
Deep Visual Forced Alignment: Learning to Align Transcription with Talking Face Video27 Feb 2023 0 repositories listed
-
17 Feb 2023 0 repositories listed
-
17 Feb 2023 0 repositories listed
-
Prompt Tuning of Deep Neural Networks for Speaker-adaptive Visual Speech Recognition16 Feb 2023 0 repositories listed
-
AV-data2vec: Self-supervised Learning of Audio-Visual Speech Representations with Contextualized Target Representations10 Feb 2023 0 repositories listed
-
A Multi-Purpose Audio-Visual Corpus for Multi-Modal Persian Speech Recognition: the Arman-AV Dataset21 Jan 2023 0 repositories listed
-
ReVISE: Self-Supervised Speech Resynthesis With Visual Input for Universal and Generalized Speech Regeneration1 Jan 2023 0 repositories listed
-
21 Dec 2022 0 repositories listed
-
Leveraging Modality-specific Representations for Audio-visual Speech Recognition via Reinforcement Learning10 Dec 2022 0 repositories listed
-
VATLM: Visual-Audio-Text Pre-Training with Unified Masked Prediction for Speech Representation Learning21 Nov 2022 0 repositories listed
-
Streaming Audio-Visual Speech Recognition with Alignment Regularization3 Nov 2022 0 repositories listed
-
29 Aug 2022 0 repositories listed
-
Kaggle Competition: Cantonese Audio-Visual Speech Recognition for In-car Commands6 Jul 2022 0 repositories listed
-
Lip-Listening: Mixing Senses to Understand Lips using Cross Modality Knowledge Distillation for Word-Based Models5 Jun 2022 0 repositories listed
-
RUSAVIC Corpus: Russian Audio-Visual Speech in Cars1 Jun 2022 0 repositories listed
-
Is Lip Region-of-Interest Sufficient for Lipreading?28 May 2022 0 repositories listed
-
Deep Learning for Visual Speech Analysis: A Survey22 May 2022 0 repositories listed
-
Learning Contextually Fused Audio-visual Representations for Audio-visual Speech Recognition15 Feb 2022 0 repositories listed
-
Transformer-Based Video Front-Ends for Audio-Visual Speech Recognition for Single and Multi-Person Video25 Jan 2022 0 repositories listed
-
Recent Progress in the CUHK Dysarthric Speech Recognition System15 Jan 2022 0 repositories listed
-
Leveraging Uni-Modal Self-Supervised Learning for Multimodal Audio-visual Speech Recognition16 Nov 2021 0 repositories listed
-
Advances and Challenges in Deep Lip Reading15 Oct 2021 0 repositories listed
-
14 Oct 2021 0 repositories listed
-
Perception Point: Identifying Critical Learning Periods in Speech for Bilingual Networks13 Oct 2021 0 repositories listed
-
Audio-Visual Speech Recognition is Worth 32×32×8 Voxels20 Sep 2021 0 repositories listed
-
LRWR: Large-Scale Benchmark for Lip Reading in Russian language14 Sep 2021 0 repositories listed
-
Large-vocabulary Audio-visual Speech Recognition in Noisy Environments10 Sep 2021 0 repositories listed
-
Spatio-Temporal Attention Mechanism and Knowledge Distillation for Lip Reading7 Aug 2021 0 repositories listed
-
Interactive decoding of words from visual speech recognition models1 Jul 2021 0 repositories listed
-
Fusing information streams in end-to-end audio-visual speech recognition19 Apr 2021 0 repositories listed
-
14 Dec 2020 0 repositories listed
-
25 Oct 2020 0 repositories listed
-
"Notic My Speech" -- Blending Speech Patterns With Multimedia12 Jun 2020 0 repositories listed
-
6 Jan 2020 0 repositories listed
-
Detecting Adversarial Attacks On Audiovisual Speech Recognition18 Dec 2019 0 repositories listed
-
Continuous Speech Recognition using EEG and Video16 Dec 2019 0 repositories listed
-
28 Nov 2019 0 repositories listed
-
Investigating the Lombard Effect Influence on End-to-End Audio-Visual Speech Recognition5 Jun 2019 0 repositories listed
-
MobiVSR: A Visual Speech Recognition Solution for Mobile Devices10 May 2019 0 repositories listed
-
End-to-End Visual Speech Recognition for Small-Scale Datasets2 Apr 2019 0 repositories listed
-
Modality Attention for End-to-End Audio-visual Speech Recognition13 Nov 2018 0 repositories listed
-
3D Feature Pyramid Attention Module for Robust Visual Speech Recognition15 Oct 2018 0 repositories listed
-
28 Sep 2018 0 repositories listed
-
Perfect match: Improved cross-modal embeddings for audio-visual synchronisation21 Sep 2018 0 repositories listed
-
13 Jul 2018 0 repositories listed
-
Deep Lip Reading: a comparison of models and an online application15 Jun 2018 0 repositories listed
-
Towards Lipreading Sentences with Active Appearance Models29 May 2018 0 repositories listed
-
Task-dependent modulation of the visual sensory thalamus assists visual-speech recognition24 May 2018 0 repositories listed
-
18 Feb 2018 0 repositories listed
-
Combining Multiple Views for Visual Speech Recognition19 Oct 2017 0 repositories listed
-
Visual Speech Recognition Using PCA Networks and LSTMs in a Tandem GMM-HMM System19 Oct 2017 0 repositories listed
-
Resolution limits on visual speech recognition3 Oct 2017 0 repositories listed
-
Visual speech recognition: aligning terminologies for better understanding3 Oct 2017 0 repositories listed
-
Which phoneme-to-viseme maps best improve visual-only computer lip-reading?3 Oct 2017 0 repositories listed
-
Multimodal Machine Learning: Integrating Language, Vision and Speech1 Jul 2017 0 repositories listed
-
Towards Estimating the Upper Bound of Visual-Speech Recognition: The Visual Lip-Reading Feasibility Database26 Apr 2017 0 repositories listed
-
Deep Multimodal Representation Learning from Temporal Data11 Apr 2017 0 repositories listed
-
End-To-End Visual Speech Recognition With LSTMs20 Jan 2017 0 repositories listed
-
Auxiliary Multimodal LSTM for Audio-visual Speech Recognition and Lipreading16 Jan 2017 0 repositories listed
-
16 Nov 2016 0 repositories listed
-
Audio Visual Speech Recognition using Deep Recurrent Neural Networks9 Nov 2016 0 repositories listed
-
A three-dimensional approach to Visual Speech Recognition using Discrete Cosine Transforms7 Sep 2016 0 repositories listed
-
Manifold-Kernels Comparison in MKPLS for Visual Speech Recognition22 Jan 2016 0 repositories listed
-
Listening With Your Eyes: Towards a Practical Visual Speech Recognition System Using Deep Boltzmann Machines1 Dec 2015 0 repositories listed
-
Video-Based Action Recognition Using Rate-Invariant Analysis of Covariance Trajectories23 Mar 2015 0 repositories listed
-
Deep Multimodal Learning for Audio-Visual Speech Recognition22 Jan 2015 0 repositories listed
-
Visual Words for Automatic Lip-Reading17 Sep 2014 0 repositories listed
-
Visual Speech Recognition3 Sep 2014 0 repositories listed
-
Recognition of Isolated Words using Zernike and MFCC features for Audio Visual Speech Recognition4 Jul 2014 0 repositories listed
-
Preliminary Test of a Real-Time, Interactive Silent Speech Interface Based on Electromagnetic Articulograph1 Jun 2014 0 repositories listed
-
Rate-Invariant Analysis of Trajectories on Riemannian Manifolds with Application in Visual Speech Recognition1 Jun 2014 0 repositories listed
-
MKPLS: Manifold Kernel Partial Least Squares for Lipreading and Speaker Identification1 Jun 2013 0 repositories listed
-
Building a synchronous corpus of acoustic and 3D facial marker data for adaptive audio-visual speech synthesis1 May 2012 0 repositories listed
-
SUTAV: A Turkish Audio-Visual Database1 May 2012 0 repositories listed
Syntology lines on 1 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.