{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/talking-face-generation-by-adversarially","title":"Talking Face Generation by Adversarially Disentangled Audio-Visual Representation","arxiv_id":"1807.07860","date":"2018-07-20","proceeding":null,"authors":["Hang Zhou","Yu Liu","Ziwei Liu","Ping Luo","Xiaogang Wang"],"abstract":"Talking face generation aims to synthesize a sequence of face images that\ncorrespond to a clip of speech. This is a challenging task because face\nappearance variation and semantics of speech are coupled together in the subtle\nmovements of the talking face regions. Existing works either construct specific\nface appearance model on specific subjects or model the transformation between\nlip motion and speech. In this work, we integrate both aspects and enable\narbitrary-subject talking face generation by learning disentangled audio-visual\nrepresentation. We find that the talking face sequence is actually a\ncomposition of both subject-related information and speech-related information.\nThese two spaces are then explicitly disentangled through a novel\nassociative-and-adversarial training process. This disentangled representation\nhas an advantage where both audio and video can serve as inputs for generation.\nExtensive experiments show that the proposed approach generates realistic\ntalking face sequences on arbitrary subjects with much clearer lip motion\npatterns than previous work. We also demonstrate the learned audio-visual\nrepresentation is extremely useful for the tasks of automatic lip reading and\naudio-video retrieval.","url_abs":"http://arxiv.org/abs/1807.07860v2","url_pdf":"http://arxiv.org/pdf/1807.07860v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"talking-face-generation-by-adversarially","repo_url":"https://github.com/Hangz-nju-cuhk/Talking-Face-Generation-DAVS","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"face-generation","task_name":"Face Generation"},{"task_slug":"lip-reading","task_name":"Lip Reading"},{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":"talking-face-generation","task_name":"Talking Face Generation"},{"task_slug":"video-retrieval","task_name":"Video Retrieval"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1807.07860","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1807.07860"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/Hangz-nju-cuhk/Talking-Face-Generation-DAVS","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"ran_draft_wrong":2,"unverified":3},"by_repo_kind":{"listed":{"samples":5,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"8c6b999f5a291101","entry":"to_np","repo":"Hangz-nju-cuhk/Talking-Face-Generation-DAVS","repo_kind":"listed","path":"embedding_utils.py","file_url":"https://github.com/Hangz-nju-cuhk/Talking-Face-Generation-DAVS/blob/HEAD/embedding_utils.py","link_basis":"harvester_set","language":"python","status":"ran_draft_wrong","verification_level":2,"contract_check":"MISDECLARED","metamorphic_tier":"invariant","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"8c6b999f5a291101"}},{"code_sha256_prefix":"5baaa8c1b148ef70","entry":"conv3x3","repo":"Hangz-nju-cuhk/Talking-Face-Generation-DAVS","repo_kind":"listed","path":"network/FAN_feature_extractor.py","file_url":"https://github.com/Hangz-nju-cuhk/Talking-Face-Generation-DAVS/blob/HEAD/network/FAN_feature_extractor.py","link_basis":"harvester_set","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"5baaa8c1b148ef70"}},{"code_sha256_prefix":"167513eefb3a1218","entry":"copy_state_dict","repo":"Hangz-nju-cuhk/Talking-Face-Generation-DAVS","repo_kind":"listed","path":"embedding_utils.py","file_url":"https://github.com/Hangz-nju-cuhk/Talking-Face-Generation-DAVS/blob/HEAD/embedding_utils.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"167513eefb3a1218"}},{"code_sha256_prefix":"a8cd382a09b3276a","entry":"load_checkpoint","repo":"Hangz-nju-cuhk/Talking-Face-Generation-DAVS","repo_kind":"listed","path":"embedding_utils.py","file_url":"https://github.com/Hangz-nju-cuhk/Talking-Face-Generation-DAVS/blob/HEAD/embedding_utils.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"a8cd382a09b3276a"}},{"code_sha256_prefix":"11a56c4af50d7c77","entry":"transformation_from_points","repo":"Hangz-nju-cuhk/Talking-Face-Generation-DAVS","repo_kind":"listed","path":"preprocess/face_align.py","file_url":"https://github.com/Hangz-nju-cuhk/Talking-Face-Generation-DAVS/blob/HEAD/preprocess/face_align.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"11a56c4af50d7c77"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}