Papers › Photographic Image Synthesis with Cascaded Refinement Networks

Photographic Image Synthesis with Cascaded Refinement Networks

28 Jul 2017ICCV 2017 10arXiv:1707.09405archive 2025-07-28

Qifeng Chen, Vladlen Koltun

We present an approach to synthesizing photographic images conditioned on semantic layouts. Given a semantic label map, our approach produces an image with photographic appearance that conforms to the input layout. The approach thus functions as a rendering engine that takes a two-dimensional semantic specification of the scene and produces a corresponding photographic image. Unlike recent and contemporaneous work, our approach does not rely on adversarial training. We show that photographic images can be synthesized from semantic layouts by a single feedforward network with appropriate structure, trained end-to-end with a direct regression objective. The presented approach scales seamlessly to high resolutions; we demonstrate this by synthesizing photographic images at 2-megapixel resolution, the full resolution of our training data. Extensive perceptual experiments on datasets of outdoor and indoor scenes demonstrate that images synthesized by the presented approach are considerably more realistic than alternative approaches. The results are shown in the supplementary video at https://youtu.be/0fhUJT21-bs

PaperPDFConference PDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Image GenerationImage-to-Image Translation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Image-to-Image Translation ADE20K Labels-to-Photos CRN Accuracy 68.8% #9 of 16 Archive leaderboard report
Image-to-Image Translation ADE20K Labels-to-Photos CRN FID 73.3 #9 of 16 Archive leaderboard report
Image-to-Image Translation ADE20K Labels-to-Photos CRN mIoU 22.4 #9 of 16 Archive leaderboard report
Image-to-Image Translation ADE20K-Outdoor Labels-to-Photos CRN Accuracy 68.6% #5 of 7 Archive leaderboard report
Image-to-Image Translation ADE20K-Outdoor Labels-to-Photos CRN FID 99.0 #5 of 7 Archive leaderboard report
Image-to-Image Translation ADE20K-Outdoor Labels-to-Photos CRN mIoU 16.5 #5 of 7 Archive leaderboard report
Image-to-Image Translation COCO-Stuff Labels-to-Photos CRN Accuracy 40.4% #14 of 15 Archive leaderboard report
Image-to-Image Translation COCO-Stuff Labels-to-Photos CRN FID 70.4 #14 of 15 Archive leaderboard report
Image-to-Image Translation COCO-Stuff Labels-to-Photos CRN mIoU 23.7 #14 of 15 Archive leaderboard report
Image-to-Image Translation Cityscapes Labels-to-Photo CRN FID 104.7 #11 of 21 Archive leaderboard report
Image-to-Image Translation Cityscapes Labels-to-Photo CRN Per-pixel Accuracy 77.1% #11 of 21 Archive leaderboard report
Image-to-Image Translation Cityscapes Labels-to-Photo CRN mIoU 52.4 #11 of 21 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Dense ConnectionsFeedforward Network

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections