Papers › Visual Speech Enhancement Without A Real Visual Stream

Visual Speech Enhancement Without A Real Visual Stream

20 Dec 2020arXiv:2012.10852archive 2025-07-28

Sindhu B Hegde, K R Prajwal, Rudrabha Mukhopadhyay, Vinay Namboodiri, C. V. Jawahar

In this work, we re-think the task of speech enhancement in unconstrained real-world environments. Current state-of-the-art methods use only the audio stream and are limited in their performance in a wide range of real-world noises. Recent works using lip movements as additional cues improve the quality of generated speech over "audio-only" methods. But, these methods cannot be used for several applications where the visual stream is unreliable or completely absent. We propose a new paradigm for speech enhancement by exploiting recent breakthroughs in speech-driven lip synthesis. Using one such model as a teacher network, we train a robust student network to produce accurate lip movements that mask away the noise, thus acting as a "visual noise filter". The intelligibility of the speech enhanced by our pseudo-lip approach is comparable (< 3% difference) to the case of using real lips. This implies that we can exploit the advantages of using lip movements even in the absence of a real video stream. We rigorously evaluate our model using quantitative metrics as well as human evaluations. Additional ablation studies and a demo video on our website containing qualitative comparisons and results clearly illustrate the effectiveness of our approach. We provide a demo video which clearly illustrates the effectiveness of our proposed approach on our website: \url{http://cvit.iiit.ac.in/research/projects/cvit-projects/visual-speech-enhancement-without-a-real-visual-stream}. The code and models are also released for future research: \url{https://github.com/Sindhu-Hegde/pseudo-visual-speech-denoising}.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

Sindhu-Hegde/pseudo-visual-speech-denoising officialmentioned in papermentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

DenoisingSpeech DenoisingSpeech Enhancement

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Speech Denoising LRS2+VGGSound CBAK 2.41 #1 of 1 Archive leaderboard report
Speech Denoising LRS2+VGGSound COVL 2.15 #1 of 1 Archive leaderboard report
Speech Denoising LRS2+VGGSound CSIG 3.16 #1 of 1 Archive leaderboard report
Speech Denoising LRS2+VGGSound PESQ 2.71 #1 of 1 Archive leaderboard report
Speech Denoising LRS2+VGGSound STOI 0.87 #1 of 1 Archive leaderboard report
Speech Denoising LRS3+VGGSound CBAK 2.47 #1 of 1 Archive leaderboard report
Speech Denoising LRS3+VGGSound COVL 2.25 #1 of 1 Archive leaderboard report
Speech Denoising LRS3+VGGSound CSIG 3.18 #1 of 1 Archive leaderboard report
Speech Denoising LRS3+VGGSound PESQ 2.72 #1 of 1 Archive leaderboard report
Speech Denoising LRS3+VGGSound STOI 0.88 #1 of 1 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections