{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/visual-object-networks-image-generation-with","title":"Visual Object Networks: Image Generation with Disentangled 3D Representation","arxiv_id":"1812.02725","date":"2018-12-06","proceeding":"NeurIPS 2018","authors":["Jun-Yan Zhu","Zhoutong Zhang","Chengkai Zhang","Jiajun Wu","Antonio Torralba","Joshua B. Tenenbaum","William T. Freeman"],"abstract":"Recent progress in deep generative models has led to tremendous breakthroughs\nin image generation. However, while existing models can synthesize\nphotorealistic images, they lack an understanding of our underlying 3D world.\nWe present a new generative model, Visual Object Networks (VON), synthesizing\nnatural images of objects with a disentangled 3D representation. Inspired by\nclassic graphics rendering pipelines, we unravel our image formation process\ninto three conditionally independent factors---shape, viewpoint, and\ntexture---and present an end-to-end adversarial learning framework that jointly\nmodels 3D shapes and 2D images. Our model first learns to synthesize 3D shapes\nthat are indistinguishable from real shapes. It then renders the object's 2.5D\nsketches (i.e., silhouette and depth map) from its shape under a sampled\nviewpoint. Finally, it learns to add realistic texture to these 2.5D sketches\nto generate natural images. The VON not only generates images that are more\nrealistic than state-of-the-art 2D image synthesis methods, but also enables\nmany 3D operations such as changing the viewpoint of a generated image, editing\nof shape and texture, linear interpolation in texture and shape space, and\ntransferring appearance across different objects and viewpoints.","url_abs":"http://arxiv.org/abs/1812.02725v1","url_pdf":"http://arxiv.org/pdf/1812.02725v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"visual-object-networks-image-generation-with","repo_url":"https://github.com/junyanz/VON","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"NOASSERTION"}}],"tasks":[{"task_slug":"image-generation","task_name":"Image Generation"},{"task_slug":"object","task_name":"Object"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1812.02725","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}