Papers › Unlocking Pre-trained Image Backbones for Semantic Image Synthesis

Unlocking Pre-trained Image Backbones for Semantic Image Synthesis

20 Dec 2023CVPR 2024 1arXiv:2312.13314archive 2025-07-28

Tariq Berrada, Jakob Verbeek, Camille Couprie, Karteek Alahari

Semantic image synthesis, i.e., generating images from user-provided semantic label maps, is an important conditional image generation task as it allows to control both the content as well as the spatial layout of generated images. Although diffusion models have pushed the state of the art in generative image modeling, the iterative nature of their inference process makes them computationally demanding. Other approaches such as GANs are more efficient as they only need a single feed-forward pass for generation, but the image quality tends to suffer on large and diverse datasets. In this work, we propose a new class of GAN discriminators for semantic image synthesis that generates highly realistic images by exploiting feature backbone networks pre-trained for tasks such as image classification. We also introduce a new generator architecture with better context modeling and using cross-attention to inject noise into latent variables, leading to more diverse generated images. Our model, which we dub DP-SIMS, achieves state-of-the-art results in terms of image quality and consistency with the input label maps on ADE-20K, COCO-Stuff, and Cityscapes, surpassing recent diffusion models while requiring two orders of magnitude less compute for inference.

PaperPDFConference PDF

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Conditional Image GenerationImage ClassificationImage GenerationImage-to-Image Translationimage-classification

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Image-to-Image Translation ADE20K Labels-to-Photos DP-SIMS (ConvNext-L) FID 22.7 #1 of 16 Archive leaderboard report
Image-to-Image Translation ADE20K Labels-to-Photos DP-SIMS (ConvNext-L) mIoU 54.3 #1 of 16 Archive leaderboard report
Image-to-Image Translation COCO-Stuff Labels-to-Photos DP-SIMS (ConvNext-XL) FID 13.3 #1 of 15 Archive leaderboard report
Image-to-Image Translation COCO-Stuff Labels-to-Photos DP-SIMS (ConvNext-L) FID 13.6 #2 of 15 Archive leaderboard report
Image-to-Image Translation Cityscapes Labels-to-Photo DP-SIMS (ConvNext-L) FID 38.2 #1 of 21 Archive leaderboard report
Image-to-Image Translation Cityscapes Labels-to-Photo DP-SIMS (ConvNext-L) mIoU 76.3 #1 of 21 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Diffusion

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections