Browse State-of-the-Art › Talking Face Generation
Talking Face Generation
43 papers with code · 2 benchmarks · 6 datasets archive 2025-07-28
Talking face generation aims to synthesize a sequence of face images that correspond to given speech semantics
( Image credit: Talking Face Generation by Adversarially Disentangled Audio-Visual Representation )
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
2 leaderboard tables shown for this task, 2 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| CREMA-D (1 row) | EmoGen | Emotionally Enhanced Talking Face Generation | code | — | Compare |
| LRW (1 row) | LipGAN | Towards Automatic Face-to-Face Translation | code | Syntology ran 2 of 2 samples · 0 unverified | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
6 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
2 subtasks in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 43 papers with code (110 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
23 Aug 2020 4 repositories listed Syntology ran 1 of 2 samples · 1 unverified · 2 pointer-only (licence)However, they fail to accurately morph the lip movements of arbitrary identities in dynamic, unconstrained talking face videos, resulting in significant parts of the video being out-of-sync with the new audio.
-
27 Apr 2020 3 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)We present a method that generates expressive talking heads from a single facial image with audio as the only input.
-
22 Nov 2022 2 repositories listed Syntology ran 9 of 11 samples · 2 unverifiedWhile dynamic Neural Radiance Fields (NeRF) have shown success in high-fidelity 3D modeling of talking portraits, the slow training and inference speed severely obstruct their potential usage.
-
3 Jan 2025 1 repository listedSignificant progress has been made in talking-face video generation research; however, precise lip-audio synchronization and high visual quality remain challenging in editing lip shapes based on input audio.
-
18 Dec 2024 1 repository listedRecent advances in co-speech gesture and talking head generation have been impressive, yet most methods focus on only one of the two tasks.
-
9 Sep 2024 1 repository listedIn this paper, we propose the KFusion of Dual-Domain model, a robust model that generates landmarks from audio.
-
5 Jun 2024 1 repository listedAudio-driven talking face generation has garnered significant interest within the domain of digital human research.
-
26 Mar 2024 1 repository listedDeepfake is a technology dedicated to creating highly realistic facial images and videos under specific conditions, which has significant application potential in fields such as entertainment, movie production, digital…
-
16 Jan 2024 1 repository listed Syntology ran 5 of 6 samples · 1 unverifiedOne-shot 3D talking portrait generation aims to reconstruct a 3D avatar from an unseen image, and then animate it with a reference video or audio to generate a talking portrait video.
-
12 Dec 2023 1 repository listedOur proposed GSmoothFace model mainly consists of the Audio to Expression Prediction (A2EP) module and the Target Adaptive Face Translation (TAFT) module.
-
11 Dec 2023 1 repository listedOur method, which we call NEUral Text to ARticulate Talk (NEUTART), is a talking face generator that uses a joint audiovisual feature space, as well as speech-informed 3D facial reconstructions and a lip-reading loss…
-
29 Nov 2023 1 repository listed Syntology ran 3 of 4 samples · 1 unverified · 4 pointer-only (licence)A lifelike talking head requires synchronized coordination of subject identity, lip movements, facial expressions, and head poses.
-
3 Nov 2023 1 repository listedOur rigorous experiments comprehensively highlight that our ground-breaking approach outpaces existing methods with considerable margins and delivers seamless, intelligible videos in person-generic and multilingual…
-
9 Oct 2023 1 repository listedFirst, FaceEncoder is used to obtain latent code by extracting features from the visual face information taken from the video source containing the face frame.
-
14 Sep 2023 1 repository listedIn particular, we propose a Fine-Grained Feature Fusion (FGFF) module to effectively capture fine texture feature information around teeth and surrounding regions, and use these features to fine-grain the feature map to…
-
15 May 2023 1 repository listed Syntology ran 4 of 6 samples · 2 unverifiedPrior landmark characteristics of the speaker's face are employed to make the generated landmarks coincide with the facial outline of the speaker.
-
29 Mar 2023 1 repository listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)To address the problem, we propose using a lip-reading expert to improve the intelligibility of the generated lip regions by penalizing the incorrect generation results.
-
21 Mar 2023 1 repository listedTo mitigate this, we build a talking face generation framework conditioned on a categorical emotion to generate videos with appropriate expressions, making them more realistic and convincing.
-
7 Mar 2023 1 repository listed Syntology ran 2 of 10 samples · 8 unverified · 10 pointer-only (licence)Different from previous works relying on multiple up-sample layers to directly generate pixels from latent embeddings, DINet performs spatial deformation on feature maps of reference images to better preserve…
-
31 Jan 2023 1 repository listed Syntology ran 3 of 4 samples · 1 unverifiedGenerating photo-realistic video portrait with arbitrary speech audio is a crucial problem in film-making and virtual reality.
-
16 Jan 2023 1 repository listed Syntology ran 15 of 23 samples · 8 unverified · 1 pointer-only (licence)In this paper, we introduce a novel self-supervised disentanglement framework to decouple pose and expression without 3DMMs and paired data, which consists of a motion editing module, a pose generator, and an expression…
-
3 Jan 2023 1 repository listed Syntology ran 0 of 2 samples · 2 unverifiedIn a nutshell, we aim to attain a speaking style from an arbitrary reference speaking video and then drive the one-shot portrait to speak with the reference speaking style and another piece of audio.
-
21 Sep 2022 1 repository listed Syntology ran 5 of 11 samples · 6 unverifiedIn FNeVR, we design a 3D Face Volume Rendering (FVR) module to enhance the facial details for image rendering.
-
24 Jul 2022 1 repository listed Syntology ran 1 of 1 samples · 0 unverifiedThus the facial radiance field can be flexibly adjusted to the new identity with few reference images.
-
1 Jun 2022 1 repository listedWe introduce the Merkel Podcast Corpus, an audio-visual-text corpus in German collected from 16 years of (almost) weekly Internet podcasts of former German chancellor Angela Merkel.
-
24 May 2022 1 repository listedWe introduce the Merkel Podcast Corpus, an audio-visual-text corpus in German collected from 16 years of (almost) weekly Internet podcasts of former German chancellor Angela Merkel.
-
8 Mar 2022 1 repository listed Syntology ran 1 of 8 samples · 7 unverifiedOur framework elevates the resolution of the synthesized talking face to 1024*1024 for the first time, even though the training dataset has a lower resolution.
-
22 Sep 2021 1 repository listedThe first stage is a deep neural network that extracts deep audio features along with a manifold projection to project the features to the target person's speech space.
-
18 Aug 2021 1 repository listed Syntology ran 3 of 3 samples · 0 unverified · 3 pointer-only (licence)In this paper, we propose a talking face generation method that takes an audio signal as input and a short target video clip as reference, and synthesizes a photo-realistic video of the target face with natural lip…
-
14 Jul 2021 1 repository listedHowever, the AR decoding manner generates current lip frame conditioned on frames generated previously, which inherently hinders the inference speed, and also has a detrimental effect on the quality of generated lip…
Syntology lines on 15 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections