{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/pix3d-dataset-and-methods-for-single-image-3d","title":"Pix3D: Dataset and Methods for Single-Image 3D Shape Modeling","arxiv_id":"1804.04610","date":"2018-04-12","proceeding":"CVPR 2018 6","authors":["Xingyuan Sun","Jiajun Wu","Xiuming Zhang","Zhoutong Zhang","Chengkai Zhang","Tianfan Xue","Joshua B. Tenenbaum","William T. Freeman"],"abstract":"We study 3D shape modeling from a single image and make contributions to it\nin three aspects. First, we present Pix3D, a large-scale benchmark of diverse\nimage-shape pairs with pixel-level 2D-3D alignment. Pix3D has wide applications\nin shape-related tasks including reconstruction, retrieval, viewpoint\nestimation, etc. Building such a large-scale dataset, however, is highly\nchallenging; existing datasets either contain only synthetic data, or lack\nprecise alignment between 2D images and 3D shapes, or only have a small number\nof images. Second, we calibrate the evaluation criteria for 3D shape\nreconstruction through behavioral studies, and use them to objectively and\nsystematically benchmark cutting-edge reconstruction algorithms on Pix3D.\nThird, we design a novel model that simultaneously performs 3D reconstruction\nand pose estimation; our multi-task learning approach achieves state-of-the-art\nperformance on both tasks.","url_abs":"http://arxiv.org/abs/1804.04610v1","url_pdf":"http://arxiv.org/pdf/1804.04610v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"pix3d-dataset-and-methods-for-single-image-3d","repo_url":"https://github.com/xingyuansun/pix3d","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"tf","reach":{"status":"ok"}}],"tasks":[{"task_slug":"3d-reconstruction","task_name":"3D Reconstruction"},{"task_slug":"3d-shape-modeling","task_name":"3D Shape Modeling"},{"task_slug":"3d-shape-reconstruction","task_name":"3D Shape Reconstruction"},{"task_slug":"multi-task-learning","task_name":"Multi-Task Learning"},{"task_slug":"pose-estimation","task_name":"Pose Estimation"},{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":"viewpoint-estimation","task_name":"Viewpoint Estimation"}],"methods":[],"datasets_introduced":[{"slug":"pix3d","name":"Pix3D","full_name":""}],"methods_introduced":[],"results":[{"leaderboard":"/sota/3d-shape-retrieval-on-pix3d","task":"3D Shape Classification","dataset":"Pix3D","model":"MarrNet extension (w/o Pose)","rank_in_archive_order":1,"of":3,"metrics":{"R@1":"0.53","R@16":"0.85","R@2":"0.62","R@32":"0.90","R@4":"0.71","R@8":"0.78"},"uses_additional_data":false},{"leaderboard":"/sota/3d-shape-reconstruction-on-pix3d","task":"3D Shape Reconstruction","dataset":"Pix3D","model":"MarrNet extension (w/ Pose)","rank_in_archive_order":4,"of":5,"metrics":{"CD":"0.119","EMD":"0.118","IoU":"0.282"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1804.04610","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}