{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/lip-movements-generation-at-a-glance","title":"Lip Movements Generation at a Glance","arxiv_id":"1803.10404","date":"2018-03-28","proceeding":"ECCV 2018 9","authors":["Lele Chen","Zhiheng Li","Ross K. Maddox","Zhiyao Duan","Chenliang Xu"],"abstract":"Cross-modality generation is an emerging topic that aims to synthesize data\nin one modality based on information in a different modality. In this paper, we\nconsider a task of such: given an arbitrary audio speech and one lip image of\narbitrary target identity, generate synthesized lip movements of the target\nidentity saying the speech. To perform well in this task, it inevitably\nrequires a model to not only consider the retention of target identity,\nphoto-realistic of synthesized images, consistency and smoothness of lip images\nin a sequence, but more importantly, learn the correlations between audio\nspeech and lip movements. To solve the collective problems, we explore the best\nmodeling of the audio-visual correlations in building and training a\nlip-movement generator network. Specifically, we devise a method to fuse audio\nand image embeddings to generate multiple lip images at once and propose a\nnovel correlation loss to synchronize lip changes and speech changes. Our final\nmodel utilizes a combination of four losses for a comprehensive consideration\nin generating lip movements; it is trained in an end-to-end fashion and is\nrobust to lip shapes, view angles and different facial characteristics.\nThoughtful experiments on three datasets ranging from lab-recorded to lips\nin-the-wild show that our model significantly outperforms other\nstate-of-the-art methods extended to this task.","url_abs":"http://arxiv.org/abs/1803.10404v3","url_pdf":"http://arxiv.org/pdf/1803.10404v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"lip-movements-generation-at-a-glance","repo_url":"https://github.com/lelechen63/3d_gan","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1803.10404","atlas_url":"https://app.syntology.ai/?focus=1803.10404","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}