{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/object-pose-estimation-from-monocular-image","title":"Object Pose Estimation from Monocular Image using Multi-View Keypoint Correspondence","arxiv_id":"1809.00553","date":"2018-09-03","proceeding":null,"authors":["Jogendra Nath Kundu","Rahul M. V.","Aditya Ganeshan","R. Venkatesh Babu"],"abstract":"Understanding the geometry and pose of objects in 2D images is a fundamental\nnecessity for a wide range of real world applications. Driven by deep neural\nnetworks, recent methods have brought significant improvements to object pose\nestimation. However, they suffer due to scarcity of keypoint/pose-annotated\nreal images and hence can not exploit the object's 3D structural information\neffectively. In this work, we propose a data-efficient method which utilizes\nthe geometric regularity of intraclass objects for pose estimation. First, we\nlearn pose-invariant local descriptors of object parts from simple 2D RGB\nimages. These descriptors, along with keypoints obtained from renders of a\nfixed 3D template model are then used to generate keypoint correspondence maps\nfor a given monocular real image. Finally, a pose estimation network predicts\n3D pose of the object using these correspondence maps. This pipeline is further\nextended to a multi-view approach, which assimilates keypoint information from\ncorrespondence sets generated from multiple views of the 3D template model.\nFusion of multi-view information significantly improves geometric comprehension\nof the system which in turn enhances the pose estimation performance.\nFurthermore, use of correspondence framework responsible for the learning of\npose invariant keypoint descriptor also allows us to effectively alleviate the\ndata-scarcity problem. This enables our method to achieve state-of-the-art\nperformance on multiple real-image viewpoint estimation datasets, such as\nPascal3D+ and ObjectNet3D. To encourage reproducible research, we have released\nthe codes for our proposed approach.","url_abs":"http://arxiv.org/abs/1809.00553v1","url_pdf":"http://arxiv.org/pdf/1809.00553v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"object-pose-estimation-from-monocular-image","repo_url":"https://github.com/val-iisc/pose_estimation","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"object-pose-estimation-from-monocular-image","repo_url":"https://github.com/val-iisc/iSPA-Net","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"pose-estimation","task_name":"Pose Estimation"},{"task_slug":"viewpoint-estimation","task_name":"Viewpoint Estimation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1809.00553","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}