Papers › TediGAN: Text-Guided Diverse Face Image Generation and Manipulation

TediGAN: Text-Guided Diverse Face Image Generation and Manipulation

6 Dec 2020CVPR 2021 1arXiv:2012.03308archive 2025-07-28

Weihao Xia, Yujiu Yang, Jing-Hao Xue, Baoyuan Wu

In this work, we propose TediGAN, a novel framework for multi-modal image generation and manipulation with textual descriptions. The proposed method consists of three components: StyleGAN inversion module, visual-linguistic similarity learning, and instance-level optimization. The inversion module maps real images to the latent space of a well-trained StyleGAN. The visual-linguistic similarity learns the text-image matching by mapping the image and text into a common embedding space. The instance-level optimization is for identity preservation in manipulation. Our model can produce diverse and high-quality images with an unprecedented resolution at 1024. Using a control mechanism based on style-mixing, our TediGAN inherently supports image synthesis with multi-modal inputs, such as sketches or semantic labels, with or without instance guidance. To facilitate text-guided multi-modal synthesis, we propose the Multi-Modal CelebA-HQ, a large-scale dataset consisting of real face images and corresponding semantic segmentation map, sketch, and textual descriptions. Extensive experiments on the introduced dataset demonstrate the superior performance of our proposed method. Code and data are available at https://github.com/weihaox/TediGAN.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

weihaox/TediGAN officialmentioned in papermentioned on GitHubpytorchMIT report
IIGROUP/Multi-Modal-CelebA-HQ-Dataset mentioned on GitHubpytorch report
IIGROUP/TediGAN mentioned on GitHubpytorch report
iigroup/mm-celeba-hq-dataset mentioned on GitHubpytorch report
weihaox/Multi-Modal-CelebA-HQ-Dataset mentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Face Sketch SynthesisImage GenerationText-to-Image Generation

Datasets

Introduced by this paper, per the archive.

Multi-Modal CelebA-HQ

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Text-to-Image Generation Multi-Modal-CelebA-HQ TediGAN-A Acc 18.4 #6 of 10 Archive leaderboard report
Text-to-Image Generation Multi-Modal-CelebA-HQ TediGAN-A FID 106.37 #6 of 10 Archive leaderboard report
Text-to-Image Generation Multi-Modal-CelebA-HQ TediGAN-A LPIPS 0.456 #6 of 10 Archive leaderboard report
Text-to-Image Generation Multi-Modal-CelebA-HQ TediGAN-A Real 22.6 #6 of 10 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Adaptive Instance NormalizationConvolutionDense ConnectionsFeedforward NetworkR1 RegularizationStyleGAN

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections