{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/bb8-a-scalable-accurate-robust-to-partial","title":"BB8: A Scalable, Accurate, Robust to Partial Occlusion Method for Predicting the 3D Poses of Challenging Objects without Using Depth","arxiv_id":"1703.10896","date":"2017-03-31","proceeding":"ICCV 2017 10","authors":["Mahdi Rad","Vincent Lepetit"],"abstract":"We introduce a novel method for 3D object detection and pose estimation from\ncolor images only. We first use segmentation to detect the objects of interest\nin 2D even in presence of partial occlusions and cluttered background. By\ncontrast with recent patch-based methods, we rely on a \"holistic\" approach: We\napply to the detected objects a Convolutional Neural Network (CNN) trained to\npredict their 3D poses in the form of 2D projections of the corners of their 3D\nbounding boxes. This, however, is not sufficient for handling objects from the\nrecent T-LESS dataset: These objects exhibit an axis of rotational symmetry,\nand the similarity of two images of such an object under two different poses\nmakes training the CNN challenging. We solve this problem by restricting the\nrange of poses used for training, and by introducing a classifier to identify\nthe range of a pose at run-time before estimating it. We also use an optional\nadditional step that refines the predicted poses. We improve the\nstate-of-the-art on the LINEMOD dataset from 73.7% to 89.3% of correctly\nregistered RGB frames. We are also the first to report results on the Occlusion\ndataset using color images only. We obtain 54% of frames passing the Pose 6D\ncriterion on average on several sequences of the T-LESS dataset, compared to\nthe 67% of the state-of-the-art on the same sequences which uses both color and\ndepth. The full approach is also scalable, as a single network can be trained\nfor multiple objects simultaneously.","url_abs":"http://arxiv.org/abs/1703.10896v2","url_pdf":"http://arxiv.org/pdf/1703.10896v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"bb8-a-scalable-accurate-robust-to-partial","repo_url":"https://github.com/Microsoft/singleshotpose","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}},{"paper_slug":"bb8-a-scalable-accurate-robust-to-partial","repo_url":"https://github.com/Yongjjun/singleshotpose","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"3d-object-detection","task_name":"3D Object Detection"},{"task_slug":"6d-pose-estimation","task_name":"6D Pose Estimation using RGB"},{"task_slug":"object-detection","task_name":"Object Detection"},{"task_slug":"pose-estimation","task_name":"Pose Estimation"},{"task_slug":"object-detection-1","task_name":"object-detection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/6d-pose-estimation-on-linemod","task":"6D Pose Estimation using RGB","dataset":"LineMOD","model":"BB8","rank_in_archive_order":19,"of":22,"metrics":{"Accuracy":"83.9%","Accuracy (ADD)":"43.6%","Mean ADD":"43.6"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1703.10896","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}