{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/context-aware-synthesis-and-placement-of","title":"Context-Aware Synthesis and Placement of Object Instances","arxiv_id":"1812.02350","date":"2018-12-06","proceeding":"NeurIPS 2018 12","authors":["Donghoon Lee","Sifei Liu","Jinwei Gu","Ming-Yu Liu","Ming-Hsuan Yang","Jan Kautz"],"abstract":"Learning to insert an object instance into an image in a semantically\ncoherent manner is a challenging and interesting problem. Solving it requires\n(a) determining a location to place an object in the scene and (b) determining\nits appearance at the location. Such an object insertion model can potentially\nfacilitate numerous image editing and scene parsing applications. In this\npaper, we propose an end-to-end trainable neural network for the task of\ninserting an object instance mask of a specified class into the semantic label\nmap of an image. Our network consists of two generative modules where one\ndetermines where the inserted object mask should be (i.e., location and scale)\nand the other determines what the object mask shape (and pose) should look\nlike. The two modules are connected together via a spatial transformation\nnetwork and jointly trained. We devise a learning procedure that leverage both\nsupervised and unsupervised data and show our model can insert an object at\ndiverse locations with various appearances. We conduct extensive experimental\nvalidations with comparisons to strong baselines to verify the effectiveness of\nthe proposed network.","url_abs":"http://arxiv.org/abs/1812.02350v2","url_pdf":"http://arxiv.org/pdf/1812.02350v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"context-aware-synthesis-and-placement-of","repo_url":"https://github.com/NVlabs/Instance_Insertion","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":null},{"paper_slug":"context-aware-synthesis-and-placement-of","repo_url":"https://github.com/bcmi/graconet-object-placement","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"object","task_name":"Object"},{"task_slug":"scene-parsing","task_name":"Scene Parsing"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1812.02350","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}