{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/instagan-instance-aware-image-to-image","title":"InstaGAN: Instance-aware Image-to-Image Translation","arxiv_id":"1812.10889","date":"2018-12-28","proceeding":null,"authors":["Sangwoo Mo","Minsu Cho","Jinwoo Shin"],"abstract":"Unsupervised image-to-image translation has gained considerable attention due\nto the recent impressive progress based on generative adversarial networks\n(GANs). However, previous methods often fail in challenging cases, in\nparticular, when an image has multiple target instances and a translation task\ninvolves significant changes in shape, e.g., translating pants to skirts in\nfashion images. To tackle the issues, we propose a novel method, coined\ninstance-aware GAN (InstaGAN), that incorporates the instance information\n(e.g., object segmentation masks) and improves multi-instance transfiguration.\nThe proposed method translates both an image and the corresponding set of\ninstance attributes while maintaining the permutation invariance property of\nthe instances. To this end, we introduce a context preserving loss that\nencourages the network to learn the identity function outside of target\ninstances. We also propose a sequential mini-batch inference/training technique\nthat handles multiple instances with a limited GPU memory and enhances the\nnetwork to generalize better for multiple instances. Our comparative evaluation\ndemonstrates the effectiveness of the proposed method on different image\ndatasets, in particular, in the aforementioned challenging cases. Code and\nresults are available in https://github.com/sangwoomo/instagan","url_abs":"http://arxiv.org/abs/1812.10889v2","url_pdf":"http://arxiv.org/pdf/1812.10889v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"instagan-instance-aware-image-to-image","repo_url":"https://github.com/sangwoomo/instagan","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok","spdx":"NOASSERTION"}}],"tasks":[{"task_slug":null,"task_name":"GPU"},{"task_slug":"image-to-image-translation","task_name":"Image-to-Image Translation"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"},{"task_slug":"translation","task_name":"Translation"},{"task_slug":"unsupervised-image-to-image-translation","task_name":"Unsupervised Image-To-Image Translation"}],"methods":[{"method_slug":"convolution","method_name":"Convolution"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/image-to-image-translation-on-object-1","task":"Image-to-Image Translation","dataset":"Object Transfiguration (sheep-to-giraffe)","model":"InstaGAN","rank_in_archive_order":1,"of":2,"metrics":{"classification score":"78.1"},"uses_additional_data":false},{"leaderboard":"/sota/image-to-image-translation-on-object-1","task":"Image-to-Image Translation","dataset":"Object Transfiguration (sheep-to-giraffe)","model":"CycleGAN","rank_in_archive_order":2,"of":2,"metrics":{"classification score":"59.4"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1812.10889","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}