Browse State-of-the-Art › Talking Head Generation
Talking Head Generation
51 papers with code · 7 benchmarks · 3 datasets archive 2025-07-28
Talking head generation is the task of generating a talking face from a set of images of a person.
( Image credit: Few-Shot Adversarial Learning of Realistic Neural Talking Head Models )
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
7 leaderboard tables shown for this task, 7 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
3 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
1 subtask in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 51 papers with code (119 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
20 May 2019 6 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)In order to create a personalized talking head model, these works require training on a large dataset of images of a single person.
-
23 Aug 2020 4 repositories listed Syntology ran 1 of 2 samples · 1 unverified · 2 pointer-only (licence)However, they fail to accurately morph the lip movements of arbitrary identities in dynamic, unconstrained talking face videos, resulting in significant parts of the video being out-of-sync with the new audio.
-
27 Apr 2020 3 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)We present a method that generates expressive talking heads from a single facial image with audio as the only input.
-
22 Nov 2022 2 repositories listedWe present SadTalker, which generates 3D motion coefficients (head pose, expression) of the 3DMM from audio and implicitly modulates a novel 3D-aware face render for talking head generation.
-
23 Jun 2025 1 repository listedTalking Head Generation (THG) has emerged as a transformative technology in computer vision, enabling the synthesis of realistic human faces synchronized with image, audio, text, or video inputs.
-
2 Jun 2025 1 repository listedAdvances in talking-head animation based on Latent Diffusion Models (LDM) enable the creation of highly realistic, synchronized videos.
-
3 Apr 2025 1 repository listedTo this end, we introduce \textbf{ACTalker}, an end-to-end video diffusion framework that supports both multi-signals control and single-signal control for talking head video generation.
-
27 Feb 2025 1 repository listedTo fully exploit the universal motion priors to learn an unseen new identity, we then present a Motion-Aligned Adaptation strategy to adaptively align the target head to the pre-trained field, and constrain a robust…
-
18 Dec 2024 1 repository listedRecent advances in co-speech gesture and talking head generation have been impressive, yet most methods focus on only one of the two tasks.
-
12 Dec 2024 1 repository listedAudio-driven talking head generation necessitates seamless integration of audio and visual data amidst the challenges posed by diverse input portraits and intricate correlations between audio and facial motions.
-
20 Nov 2024 1 repository listedThis paper examines the integration of real-time talking-head generation for interviewer training, focusing on overcoming challenges in Audio Feature Extraction (AFE), which often introduces latency and limits…
-
17 Oct 2024 1 repository listedTo address these challenges, we present DAWN (Dynamic frame Avatar With Non-autoregressive diffusion), a framework that enables all-at-once generation of dynamic-length video sequences.
-
14 Oct 2024 1 repository listedOur extensive evaluation shows our approach performs favorably compared to fixed topology techniques, setting a new benchmark by offering a versatile and high-fidelity solution for 3D talking head generation.
-
11 Sep 2024 1 repository listedIn response to this challenge, we propose EMOdiffhead, a novel method for emotional talking head video generation that not only enables fine-grained control of emotion categories and intensities but also enables…
-
28 Mar 2024 1 repository listedAToM excels in capturing subtle lip movements by leveraging an audio attention mechanism.
-
23 Mar 2024 1 repository listedIn this work, we propose an adaptive high-quality talking-head video generation method, which synthesizes high-resolution video without additional pre-trained modules.
-
19 Mar 2024 1 repository listedThe domain of 3D talking head generation has witnessed significant progress in recent years.
-
11 Mar 2024 1 repository listedHowever, performance evaluation research lags behind the development of talking head generation techniques.
-
15 Dec 2023 1 repository listed Syntology ran 6 of 6 samples · 0 unverifiedTo more conveniently specify personalized emotions, a diffusion-based style predictor is utilized to predict the personalized emotion directly from the audio, eliminating the need for extra emotion reference.
-
29 Nov 2023 1 repository listed Syntology ran 3 of 4 samples · 1 unverified · 4 pointer-only (licence)A lifelike talking head requires synchronized coordination of subject identity, lip movements, facial expressions, and head poses.
-
3 Nov 2023 1 repository listedOur rigorous experiments comprehensively highlight that our ground-breaking approach outpaces existing methods with considerable margins and delivers seamless, intelligible videos in person-generic and multilingual…
-
10 Sep 2023 1 repository listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)Audio-driven talking-head synthesis is a popular research topic for virtual human-related applications.
-
12 Aug 2023 1 repository listedIn the second stage, an audio-driven talking head generation method is employed to produce compelling videos privided the audio generated in the first stage.
-
19 Jul 2023 1 repository listed Syntology ran 3 of 3 samples · 0 unverified · 3 pointer-only (licence)Talking head video generation aims to animate a human face in a still image with dynamic poses and expressions using motion information derived from a target-driving video, while maintaining the person's identity in the…
-
4 Jul 2023 1 repository listedAnimating still face images with deep generative models using a speech input signal is an active research topic and has seen important recent progress.
-
2 Jun 2023 1 repository listedThis paper presents a novel approach for generating 3D talking heads from raw audio inputs.
-
22 May 2023 1 repository listedIt is a large-scale digital library for head avatars with three key attributes: 1) High Fidelity: all subjects are captured by 60 synchronized, high-resolution 2K cameras in 360 degrees.
-
10 May 2023 1 repository listedIn this work, firstly, we present a novel self-supervised method for learning dense 3D facial geometry (ie, depth) from face videos, without requiring camera parameters and 3D geometry annotations in training.
-
6 Apr 2023 1 repository listed Syntology ran 3 of 12 samples · 9 unverifiedFace animation has achieved much progress in computer vision.
-
21 Mar 2023 1 repository listedTo mitigate this, we build a talking face generation framework conditioned on a categorical emotion to generate videos with appropriate expressions, making them more realistic and convincing.
Syntology lines on 8 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections