{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/monocular-object-orientation-estimation-using","title":"Monocular Object Orientation Estimation using Riemannian Regression and Classification Networks","arxiv_id":"1807.07226","date":"2018-07-19","proceeding":null,"authors":["Siddharth Mahendran","Ming Yang Lu","Haider Ali","René Vidal"],"abstract":"We consider the task of estimating the 3D orientation of an object of known\ncategory given an image of the object and a bounding box around it. Recently,\nCNN-based regression and classification methods have shown significant\nperformance improvements for this task. This paper proposes a new CNN-based\napproach to monocular orientation estimation that advances the state of the art\nin four different directions. First, we take into account the Riemannian\nstructure of the orientation space when designing regression losses and\nnonlinear activation functions. Second, we propose a mixed Riemannian\nregression and classification framework that better handles the challenging\ncase of nearly symmetric objects. Third, we propose a data augmentation\nstrategy that is specifically designed to capture changes in 3D orientation.\nFourth, our approach leads to state-of-the-art results on the PASCAL3D+\ndataset.","url_abs":"http://arxiv.org/abs/1807.07226v1","url_pdf":"http://arxiv.org/pdf/1807.07226v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"monocular-object-orientation-estimation-using","repo_url":"https://github.com/JHUVisionLab/multi-modal-regression","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"classification-1","task_name":"Classification"},{"task_slug":"data-augmentation","task_name":"Data Augmentation"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"regression-1","task_name":"regression"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}