{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/ispa-net-iterative-semantic-pose-alignment","title":"iSPA-Net: Iterative Semantic Pose Alignment Network","arxiv_id":"1808.01134","date":"2018-08-03","proceeding":null,"authors":["Jogendra Nath Kundu","Aditya Ganeshan","Rahul M. V.","Aditya Prakash","R. Venkatesh Babu"],"abstract":"Understanding and extracting 3D information of objects from monocular 2D\nimages is a fundamental problem in computer vision. In the task of 3D object\npose estimation, recent data driven deep neural network based approaches suffer\nfrom scarcity of real images with 3D keypoint and pose annotations. Drawing\ninspiration from human cognition, where the annotators use a 3D CAD model as\nstructural reference to acquire ground-truth viewpoints for real images; we\npropose an iterative Semantic Pose Alignment Network, called iSPA-Net. Our\napproach focuses on exploiting semantic 3D structural regularity to solve the\ntask of fine-grained pose estimation by predicting viewpoint difference between\na given pair of images. Such image comparison based approach also alleviates\nthe problem of data scarcity and hence enhances scalability of the proposed\napproach for novel object categories with minimal annotation. The fine-grained\nobject pose estimator is also aided by correspondence of learned spatial\ndescriptor of the input image pair. The proposed pose alignment framework\nenjoys the faculty to refine its initial pose estimation in consecutive\niterations by utilizing an online rendering setup along with effectiveness of a\nnon-uniform bin classification of pose-difference. This enables iSPA-Net to\nachieve state-of-the-art performance on various real image viewpoint estimation\ndatasets. Further, we demonstrate effectiveness of the approach for multiple\napplications. First, we show results for active object viewpoint localization\nto capture images from similar pose considering only a single image as pose\nreference. Second, we demonstrate the ability of the learned semantic\ncorrespondence to perform unsupervised part-segmentation transfer using only a\nsingle part-annotated 3D template model per object class. To encourage\nreproducible research, we have released the codes for our proposed algorithm.","url_abs":"http://arxiv.org/abs/1808.01134v1","url_pdf":"http://arxiv.org/pdf/1808.01134v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"ispa-net-iterative-semantic-pose-alignment","repo_url":"https://github.com/val-iisc/iSPA-Net","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"ispa-net-iterative-semantic-pose-alignment","repo_url":"https://github.com/val-iisc/pose_estimation","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"object","task_name":"Object"},{"task_slug":"pose-estimation","task_name":"Pose Estimation"},{"task_slug":"semantic-correspondence","task_name":"Semantic correspondence"},{"task_slug":"viewpoint-estimation","task_name":"Viewpoint Estimation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}