{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/audio2head-audio-driven-one-shot-talking-head","title":"Audio2Head: Audio-driven One-shot Talking-head Generation with Natural Head Motion","arxiv_id":"2107.09293","date":"2021-07-20","proceeding":null,"authors":["Suzhen Wang","Lincheng Li","Yu Ding","Changjie Fan","Xin Yu"],"abstract":"We propose an audio-driven talking-head method to generate photo-realistic talking-head videos from a single reference image. In this work, we tackle two key challenges: (i) producing natural head motions that match speech prosody, and (ii) maintaining the appearance of a speaker in a large head motion while stabilizing the non-face regions. We first design a head pose predictor by modeling rigid 6D head movements with a motion-aware recurrent neural network (RNN). In this way, the predicted head poses act as the low-frequency holistic movements of a talking head, thus allowing our latter network to focus on detailed facial movement generation. To depict the entire image motions arising from audio, we exploit a keypoint based dense motion field representation. Then, we develop a motion field generator to produce the dense motion fields from input audio, head poses, and a reference image. As this keypoint based representation models the motions of facial regions, head, and backgrounds integrally, our method can better constrain the spatial and temporal consistency of the generated videos. Finally, an image generation network is employed to render photo-realistic talking-head videos from the estimated keypoint based motion fields and the input reference image. Extensive experiments demonstrate that our method produces videos with plausible head motions, synchronized facial expressions, and stable backgrounds and outperforms the state-of-the-art.","url_abs":"https://arxiv.org/abs/2107.09293v1","url_pdf":"https://arxiv.org/pdf/2107.09293v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"audio2head-audio-driven-one-shot-talking-head","repo_url":"https://github.com/wangsuzhen/Audio2Head","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"image-generation","task_name":"Image Generation"},{"task_slug":"talking-head-generation","task_name":"Talking Head Generation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2107.09293","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2107.09293"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/wangsuzhen/Audio2Head","reach":null}],"summary":{"unverified":5},"by_repo_kind":{"official":{"samples":5,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":5,"samples":[{"code_sha256_prefix":"553ccb718f83ba5f","entry":"MyResNet34","repo":"wangsuzhen/Audio2Head","repo_kind":"official","path":"modules/audio2pose.py","file_url":"https://github.com/wangsuzhen/Audio2Head/blob/HEAD/modules/audio2pose.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"553ccb718f83ba5f"}},{"code_sha256_prefix":"2889fa806b425f05","entry":"ResNet","repo":"wangsuzhen/Audio2Head","repo_kind":"official","path":"modules/audio2pose.py","file_url":"https://github.com/wangsuzhen/Audio2Head/blob/HEAD/modules/audio2pose.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"2889fa806b425f05"}},{"code_sha256_prefix":"a6c2491a7ebb3c14","entry":"_resnet","repo":"wangsuzhen/Audio2Head","repo_kind":"official","path":"modules/audio2pose.py","file_url":"https://github.com/wangsuzhen/Audio2Head/blob/HEAD/modules/audio2pose.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"a6c2491a7ebb3c14"}},{"code_sha256_prefix":"a3de16936c7d63e2","entry":"audio2poseLSTM","repo":"wangsuzhen/Audio2Head","repo_kind":"official","path":"modules/audio2pose.py","file_url":"https://github.com/wangsuzhen/Audio2Head/blob/HEAD/modules/audio2pose.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"a3de16936c7d63e2"}},{"code_sha256_prefix":"240e1f9382d22bd7","entry":"resnet34","repo":"wangsuzhen/Audio2Head","repo_kind":"official","path":"modules/audio2pose.py","file_url":"https://github.com/wangsuzhen/Audio2Head/blob/HEAD/modules/audio2pose.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"240e1f9382d22bd7"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}