{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/photographic-image-synthesis-with-cascaded","title":"Photographic Image Synthesis with Cascaded Refinement Networks","arxiv_id":"1707.09405","date":"2017-07-28","proceeding":"ICCV 2017 10","authors":["Qifeng Chen","Vladlen Koltun"],"abstract":"We present an approach to synthesizing photographic images conditioned on\nsemantic layouts. Given a semantic label map, our approach produces an image\nwith photographic appearance that conforms to the input layout. The approach\nthus functions as a rendering engine that takes a two-dimensional semantic\nspecification of the scene and produces a corresponding photographic image.\nUnlike recent and contemporaneous work, our approach does not rely on\nadversarial training. We show that photographic images can be synthesized from\nsemantic layouts by a single feedforward network with appropriate structure,\ntrained end-to-end with a direct regression objective. The presented approach\nscales seamlessly to high resolutions; we demonstrate this by synthesizing\nphotographic images at 2-megapixel resolution, the full resolution of our\ntraining data. Extensive perceptual experiments on datasets of outdoor and\nindoor scenes demonstrate that images synthesized by the presented approach are\nconsiderably more realistic than alternative approaches. The results are shown\nin the supplementary video at https://youtu.be/0fhUJT21-bs","url_abs":"http://arxiv.org/abs/1707.09405v1","url_pdf":"http://arxiv.org/pdf/1707.09405v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"image-generation","task_name":"Image Generation"},{"task_slug":"image-to-image-translation","task_name":"Image-to-Image Translation"}],"methods":[{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"feedforward-network","method_name":"Feedforward Network"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/image-to-image-translation-on-ade20k-labels","task":"Image-to-Image Translation","dataset":"ADE20K Labels-to-Photos","model":"CRN","rank_in_archive_order":9,"of":16,"metrics":{"Accuracy":"68.8%","FID":"73.3","mIoU":"22.4"},"uses_additional_data":false},{"leaderboard":"/sota/image-to-image-translation-on-ade20k-outdoor","task":"Image-to-Image Translation","dataset":"ADE20K-Outdoor Labels-to-Photos","model":"CRN","rank_in_archive_order":5,"of":7,"metrics":{"Accuracy":"68.6%","FID":"99.0","mIoU":"16.5"},"uses_additional_data":false},{"leaderboard":"/sota/image-to-image-translation-on-coco-stuff","task":"Image-to-Image Translation","dataset":"COCO-Stuff Labels-to-Photos","model":"CRN","rank_in_archive_order":14,"of":15,"metrics":{"Accuracy":"40.4%","FID":"70.4","mIoU":"23.7"},"uses_additional_data":false},{"leaderboard":"/sota/image-to-image-translation-on-cityscapes","task":"Image-to-Image Translation","dataset":"Cityscapes Labels-to-Photo","model":"CRN","rank_in_archive_order":11,"of":21,"metrics":{"FID":"104.7","Per-pixel Accuracy":"77.1%","mIoU":"52.4"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1707.09405","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}