{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/end-to-end-speech-driven-facial-animation","title":"End-to-End Speech-Driven Facial Animation with Temporal GANs","arxiv_id":"1805.09313","date":"2018-05-23","proceeding":null,"authors":["Konstantinos Vougioukas","Stavros Petridis","Maja Pantic"],"abstract":"Speech-driven facial animation is the process which uses speech signals to\nautomatically synthesize a talking character. The majority of work in this\ndomain creates a mapping from audio features to visual features. This often\nrequires post-processing using computer graphics techniques to produce\nrealistic albeit subject dependent results. We present a system for generating\nvideos of a talking head, using a still image of a person and an audio clip\ncontaining speech, that doesn't rely on any handcrafted intermediate features.\nTo the best of our knowledge, this is the first method capable of generating\nsubject independent realistic videos directly from raw audio. Our method can\ngenerate videos which have (a) lip movements that are in sync with the audio\nand (b) natural facial expressions such as blinks and eyebrow movements. We\nachieve this by using a temporal GAN with 2 discriminators, which are capable\nof capturing different aspects of the video. The effect of each component in\nour system is quantified through an ablation study. The generated videos are\nevaluated based on their sharpness, reconstruction quality, and lip-reading\naccuracy. Finally, a user study is conducted, confirming that temporal GANs\nlead to more natural sequences than a static GAN-based approach.","url_abs":"http://arxiv.org/abs/1805.09313v4","url_pdf":"http://arxiv.org/pdf/1805.09313v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"end-to-end-speech-driven-facial-animation","repo_url":"https://github.com/PrashanthaTP/wav2mov","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"lip-reading","task_name":"Lip Reading"}],"methods":[{"method_slug":"convolution","method_name":"Convolution"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1805.09313","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1805.09313"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/PrashanthaTP/wav2mov","reach":null}],"summary":{"ran_violates":2},"by_repo_kind":{"listed":{"samples":2,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":2,"samples":[{"code_sha256_prefix":"c1cc5ac84d39c69f","entry":"create_batch","repo":"PrashanthaTP/wav2mov","repo_kind":"listed","path":"wav2mov/inference/generate.py","file_url":"https://github.com/PrashanthaTP/wav2mov/blob/HEAD/wav2mov/inference/generate.py","link_basis":"first_harvest_node","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"c1cc5ac84d39c69f"}},{"code_sha256_prefix":"e72a519273aeb2c6","entry":"is_exists","repo":"PrashanthaTP/wav2mov","repo_kind":"listed","path":"wav2mov/inference/generate.py","file_url":"https://github.com/PrashanthaTP/wav2mov/blob/HEAD/wav2mov/inference/generate.py","link_basis":"first_harvest_node","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"e72a519273aeb2c6"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}