Papers › Learning to Predict Layout-to-image Conditional Convolutions for Semantic Image Synthesis

Learning to Predict Layout-to-image Conditional Convolutions for Semantic Image Synthesis

15 Oct 2019NeurIPS 2019 12arXiv:1910.06809archive 2025-07-28

Xihui Liu, Guojun Yin, Jing Shao, Xiaogang Wang, Hongsheng Li

Semantic image synthesis aims at generating photorealistic images from semantic layouts. Previous approaches with conditional generative adversarial networks (GAN) show state-of-the-art performance on this task, which either feed the semantic label maps as inputs to the generator, or use them to modulate the activations in normalization layers via affine transformations. We argue that convolutional kernels in the generator should be aware of the distinct semantic labels at different locations when generating images. In order to better exploit the semantic layout for the image generator, we propose to predict convolutional kernels conditioned on the semantic label map to generate the intermediate feature maps from the noise maps and eventually generate the images. Moreover, we propose a feature pyramid semantics-embedding discriminator, which is more effective in enhancing fine details and semantic alignments between the generated images and the input semantic layouts than previous multi-scale discriminators. We achieve state-of-the-art results on both quantitative metrics and subjective evaluation on various semantic segmentation datasets, demonstrating the effectiveness of our approach.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

xh-liu/CC-FPSE officialmentioned in papermentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Image GenerationImage-to-Image TranslationSemantic Segmentation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Image-to-Image Translation ADE20K Labels-to-Photos CC-FPSE Accuracy 82.9% #7 of 16 Archive leaderboard report
Image-to-Image Translation ADE20K Labels-to-Photos CC-FPSE FID 31.7 #7 of 16 Archive leaderboard report
Image-to-Image Translation ADE20K Labels-to-Photos CC-FPSE LPIPS 0.098 #7 of 16 Archive leaderboard report
Image-to-Image Translation ADE20K Labels-to-Photos CC-FPSE mIoU 43.7 #7 of 16 Archive leaderboard report
Image-to-Image Translation COCO-Stuff Labels-to-Photos CC-FPSE Accuracy 70.7% #8 of 15 Archive leaderboard report
Image-to-Image Translation COCO-Stuff Labels-to-Photos CC-FPSE FID 19.2 #8 of 15 Archive leaderboard report
Image-to-Image Translation COCO-Stuff Labels-to-Photos CC-FPSE mIoU 41.6 #8 of 15 Archive leaderboard report
Image-to-Image Translation Cityscapes Labels-to-Photo CC-FPSE FID 54.3 #5 of 21 Archive leaderboard report
Image-to-Image Translation Cityscapes Labels-to-Photo CC-FPSE LPIPS 0.073 #5 of 21 Archive leaderboard report
Image-to-Image Translation Cityscapes Labels-to-Photo CC-FPSE Per-pixel Accuracy 82.3% #5 of 21 Archive leaderboard report
Image-to-Image Translation Cityscapes Labels-to-Photo CC-FPSE mIoU 65.5 #5 of 21 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections