{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/wavelet-based-unsupervised-label-to-image-1","title":"Wavelet-based Unsupervised Label-to-Image Translation","arxiv_id":"2305.09647","date":"2023-05-16","proceeding":"IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2022 5","authors":["George Eskandar","Mohamed Abdelsamad","Karim Armanious","Shuai Zhang","Bin Yang"],"abstract":"Semantic Image Synthesis (SIS) is a subclass of image-to-image translation where a semantic layout is used to generate a photorealistic image. State-of-the-art conditional Generative Adversarial Networks (GANs) need a huge amount of paired data to accomplish this task while generic unpaired image-to-image translation frameworks underperform in comparison, because they color-code semantic layouts and learn correspondences in appearance instead of semantic content. Starting from the assumption that a high quality generated image should be segmented back to its semantic layout, we propose a new Unsupervised paradigm for SIS (USIS) that makes use of a self-supervised segmentation loss and whole image wavelet based discrimination. Furthermore, in order to match the high-frequency distribution of real images, a novel generator architecture in the wavelet domain is proposed. We test our methodology on 3 challenging datasets and demonstrate its ability to bridge the performance gap between paired and unpaired models.","url_abs":"https://arxiv.org/abs/2305.09647v1","url_pdf":"https://arxiv.org/pdf/2305.09647v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"wavelet-based-unsupervised-label-to-image-1","repo_url":"https://github.com/GeorgeEskandar/USIS-Unsupervised-Semantic-Image-Synthesis","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"image-generation","task_name":"Image Generation"},{"task_slug":"image-to-image-translation","task_name":"Image-to-Image Translation"},{"task_slug":"multimodal-unsupervised-image-to-image","task_name":"Multimodal Unsupervised Image-To-Image Translation"},{"task_slug":"translation","task_name":"Translation"},{"task_slug":"unsupervised-image-to-image-translation","task_name":"Unsupervised Image-To-Image Translation"}],"methods":[{"method_slug":"test","method_name":"Test"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/image-to-image-translation-on-ade20k-labels","task":"Image-to-Image Translation","dataset":"ADE20K Labels-to-Photos","model":"USIS-Wavelet","rank_in_archive_order":12,"of":16,"metrics":{"FID":"34.5","mIoU":"16.95"},"uses_additional_data":false},{"leaderboard":"/sota/image-to-image-translation-on-coco-stuff","task":"Image-to-Image Translation","dataset":"COCO-Stuff Labels-to-Photos","model":"USIS-Wavelet","rank_in_archive_order":12,"of":15,"metrics":{"FID":"28.6","mIoU":"13.4"},"uses_additional_data":false},{"leaderboard":"/sota/image-to-image-translation-on-cityscapes","task":"Image-to-Image Translation","dataset":"Cityscapes Labels-to-Photo","model":"USIS-Wavelet","rank_in_archive_order":14,"of":21,"metrics":{"FID":"50.14","mIoU":"42.32"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}