Papers › Unified Generative Adversarial Networks for Controllable Image-to-Image Translation

Unified Generative Adversarial Networks for Controllable Image-to-Image Translation

12 Dec 2019arXiv:1912.06112archive 2025-07-28

Hao Tang, Hong Liu, Nicu Sebe

We propose a unified Generative Adversarial Network (GAN) for controllable image-to-image translation, i.e., transferring an image from a source to a target domain guided by controllable structures. In addition to conditioning on a reference image, we show how the model can generate images conditioned on controllable structures, e.g., class labels, object keypoints, human skeletons, and scene semantic maps. The proposed model consists of a single generator and a discriminator taking a conditional image and the target controllable structure as input. In this way, the conditional image can provide appearance information and the controllable structure can provide the structure information for generating the target result. Moreover, our model learns the image-to-image mapping through three novel losses, i.e., color loss, controllable structure guided cycle-consistency loss, and controllable structure guided self-content preserving loss. Also, we present the Fr\'echet ResNet Distance (FRD) to evaluate the quality of the generated images. Experiments on two challenging image translation tasks, i.e., hand gesture-to-gesture translation and cross-view image translation, show that our model generates convincing results, and significantly outperforms other state-of-the-art methods on both tasks. Meanwhile, the proposed framework is a unified solution, thus it can be applied to solving other controllable structure guided image translation tasks such as landmark guided facial expression translation and keypoint guided person image generation. To the best of our knowledge, we are the first to make one GAN framework work on all such controllable structure guided image translation tasks. Code is available at https://github.com/Ha0Tang/GestureGAN.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

Ha0Tang/GestureGAN officialmentioned in papermentioned on GitHubpytorchNOASSERTION report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Facial Expression TranslationGesture-to-Gesture TranslationImage GenerationImage-to-Image TranslationTranslation

1 archive task tag without a task page not shown.

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Cross-View Image-to-Image Translation Dayton (256×256) - aerial-to-ground UniGAN KL 5.17 #6 of 6 Archive leaderboard report
Cross-View Image-to-Image Translation Dayton (256×256) - aerial-to-ground UniGAN PSNR 22.0273 #6 of 6 Archive leaderboard report
Cross-View Image-to-Image Translation Dayton (256×256) - aerial-to-ground UniGAN SD 17.6542 #6 of 6 Archive leaderboard report
Cross-View Image-to-Image Translation Dayton (256×256) - aerial-to-ground UniGAN SSIM 0.3357 #6 of 6 Archive leaderboard report
Cross-View Image-to-Image Translation Dayton (64x64) - ground-to-aerial UniGAN LPIPS 0.4527 #5 of 5 Archive leaderboard report
Cross-View Image-to-Image Translation Dayton (64×64) - aerial-to-ground UniGAN KL 2.16 #3 of 5 Archive leaderboard report
Cross-View Image-to-Image Translation Dayton (64×64) - aerial-to-ground UniGAN LPIPS 0.3817 #3 of 5 Archive leaderboard report
Cross-View Image-to-Image Translation Dayton (64×64) - aerial-to-ground UniGAN PSNR 23.3632 #3 of 5 Archive leaderboard report
Cross-View Image-to-Image Translation Dayton (64×64) - aerial-to-ground UniGAN SD 16.4788 #3 of 5 Archive leaderboard report
Cross-View Image-to-Image Translation Dayton (64×64) - aerial-to-ground UniGAN SSIM 0.5064 #3 of 5 Archive leaderboard report
Cross-View Image-to-Image Translation cvusa UniGAN KL 2.6 #1 of 7 Archive leaderboard report
Cross-View Image-to-Image Translation cvusa UniGAN PSNR 22.8223 #1 of 7 Archive leaderboard report
Cross-View Image-to-Image Translation cvusa UniGAN SD 19.8276 #1 of 7 Archive leaderboard report
Cross-View Image-to-Image Translation cvusa UniGAN SSIM 0.5366 #1 of 7 Archive leaderboard report
Gesture-to-Gesture Translation NTU Hand Digit UniGAN AMT 29.3 #6 of 6 Archive leaderboard report
Gesture-to-Gesture Translation NTU Hand Digit UniGAN FID 6.7493 #6 of 6 Archive leaderboard report
Gesture-to-Gesture Translation NTU Hand Digit UniGAN FRD 1.7401 #6 of 6 Archive leaderboard report
Gesture-to-Gesture Translation NTU Hand Digit UniGAN IS 2.3783 #6 of 6 Archive leaderboard report
Gesture-to-Gesture Translation NTU Hand Digit UniGAN PSNR 32.6574 #6 of 6 Archive leaderboard report
Gesture-to-Gesture Translation Senz3D UniGAN AMT 27.6 #6 of 6 Archive leaderboard report
Gesture-to-Gesture Translation Senz3D UniGAN FID 12.4465 #6 of 6 Archive leaderboard report
Gesture-to-Gesture Translation Senz3D UniGAN FRD 2.2104 #6 of 6 Archive leaderboard report
Gesture-to-Gesture Translation Senz3D UniGAN IS 2.2159 #6 of 6 Archive leaderboard report
Gesture-to-Gesture Translation Senz3D UniGAN PSNR 31.542 #6 of 6 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

1x1 ConvolutionAverage PoolingBatch NormalizationBottleneck Residual BlockConvolutionGlobal Average PoolingKaiming InitializationMax PoolingReLUResidual BlockResidual Connection

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections