{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/visual-speech-enhancement-without-a-real","title":"Visual Speech Enhancement Without A Real Visual Stream","arxiv_id":"2012.10852","date":"2020-12-20","proceeding":null,"authors":["Sindhu B Hegde","K R Prajwal","Rudrabha Mukhopadhyay","Vinay Namboodiri","C. V. Jawahar"],"abstract":"In this work, we re-think the task of speech enhancement in unconstrained real-world environments. Current state-of-the-art methods use only the audio stream and are limited in their performance in a wide range of real-world noises. Recent works using lip movements as additional cues improve the quality of generated speech over \"audio-only\" methods. But, these methods cannot be used for several applications where the visual stream is unreliable or completely absent. We propose a new paradigm for speech enhancement by exploiting recent breakthroughs in speech-driven lip synthesis. Using one such model as a teacher network, we train a robust student network to produce accurate lip movements that mask away the noise, thus acting as a \"visual noise filter\". The intelligibility of the speech enhanced by our pseudo-lip approach is comparable (< 3% difference) to the case of using real lips. This implies that we can exploit the advantages of using lip movements even in the absence of a real video stream. We rigorously evaluate our model using quantitative metrics as well as human evaluations. Additional ablation studies and a demo video on our website containing qualitative comparisons and results clearly illustrate the effectiveness of our approach. We provide a demo video which clearly illustrates the effectiveness of our proposed approach on our website: \\url{http://cvit.iiit.ac.in/research/projects/cvit-projects/visual-speech-enhancement-without-a-real-visual-stream}. The code and models are also released for future research: \\url{https://github.com/Sindhu-Hegde/pseudo-visual-speech-denoising}.","url_abs":"https://arxiv.org/abs/2012.10852v1","url_pdf":"https://arxiv.org/pdf/2012.10852v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"visual-speech-enhancement-without-a-real","repo_url":"https://github.com/Sindhu-Hegde/pseudo-visual-speech-denoising","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"denoising","task_name":"Denoising"},{"task_slug":"speech-denoising","task_name":"Speech Denoising"},{"task_slug":"speech-enhancement","task_name":"Speech Enhancement"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/speech-denoising-on-lrs2-vggsound","task":"Speech Denoising","dataset":"LRS2+VGGSound","model":"","rank_in_archive_order":1,"of":1,"metrics":{"CBAK":"2.41","COVL":"2.15","CSIG":"3.16","PESQ":"2.71","STOI":"0.87"},"uses_additional_data":false},{"leaderboard":"/sota/speech-denoising-on-lrs3-vggsound","task":"Speech Denoising","dataset":"LRS3+VGGSound","model":"","rank_in_archive_order":1,"of":1,"metrics":{"CBAK":"2.47","COVL":"2.25","CSIG":"3.18","PESQ":"2.72","STOI":"0.88"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2012.10852","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}