{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/towards-automatic-face-to-face-translation-1","title":"Towards Automatic Face-to-Face Translation","arxiv_id":"2003.00418","date":"2020-03-01","proceeding":"ACM Multimedia, 2019 2019 10","authors":["Prajwal K R","Rudrabha Mukhopadhyay","Jerin Philip","Abhishek Jha","Vinay Namboodiri","C. V. Jawahar"],"abstract":"In light of the recent breakthroughs in automatic machine translation systems, we propose a novel approach that we term as \"Face-to-Face Translation\". As today's digital communication becomes increasingly visual, we argue that there is a need for systems that can automatically translate a video of a person speaking in language A into a target language B with realistic lip synchronization. In this work, we create an automatic pipeline for this problem and demonstrate its impact on multiple real-world applications. First, we build a working speech-to-speech translation system by bringing together multiple existing modules from speech and language. We then move towards \"Face-to-Face Translation\" by incorporating a novel visual module, LipGAN for generating realistic talking faces from the translated audio. Quantitative evaluation of LipGAN on the standard LRW test set shows that it significantly outperforms existing approaches across all standard metrics. We also subject our Face-to-Face Translation pipeline, to multiple human evaluations and show that it can significantly improve the overall user experience for consuming and interacting with multimodal content across languages. Code, models and demo video are made publicly available. Demo video: https://www.youtube.com/watch?v=aHG6Oei8jF0 Code and models: https://github.com/Rudrabha/LipGAN","url_abs":"https://arxiv.org/abs/2003.00418v1","url_pdf":"https://arxiv.org/pdf/2003.00418v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"towards-automatic-face-to-face-translation-1","repo_url":"https://github.com/Rudrabha/LipGAN","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"face-to-face-translation","task_name":"Face to Face Translation"},{"task_slug":"machine-translation","task_name":"Machine Translation"},{"task_slug":"speech-to-speech-translation","task_name":"Speech-to-Speech Translation"},{"task_slug":"translation","task_name":"Translation"},{"task_slug":"lip-sync","task_name":"Unconstrained Lip-synchronization"}],"methods":[{"method_slug":"lipgan","method_name":"LipGAN"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/talking-face-generation-on-lrw","task":"Talking Face Generation","dataset":"LRW","model":"LipGAN","rank_in_archive_order":1,"of":1,"metrics":{"LMD":"0.60","SSIM":"0.96"},"uses_additional_data":true}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2003.00418","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2003.00418"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/Rudrabha/LipGAN","reach":null}],"summary":{"ran_draft_wrong":1,"ran_honours":1},"by_repo_kind":{"official":{"samples":2,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"bec12fd12d256270","entry":"calcMaxArea","repo":"Rudrabha/LipGAN","repo_kind":"official","path":"batch_inference.py","file_url":"https://github.com/Rudrabha/LipGAN/blob/HEAD/batch_inference.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"bec12fd12d256270"}},{"code_sha256_prefix":"1040c33d9488839e","entry":"rect_to_bb","repo":"Rudrabha/LipGAN","repo_kind":"official","path":"batch_inference.py","file_url":"https://github.com/Rudrabha/LipGAN/blob/HEAD/batch_inference.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"well_formed","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"1040c33d9488839e"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}