{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-self-supervised-learning-system-for-object","title":"A Self-supervised Learning System for Object Detection using Physics Simulation and Multi-view Pose Estimation","arxiv_id":"1703.03347","date":"2017-03-09","proceeding":null,"authors":["Chaitanya Mitash","Kostas E. Bekris","Abdeslam Boularias"],"abstract":"Progress has been achieved recently in object detection given advancements in\ndeep learning. Nevertheless, such tools typically require a large amount of\ntraining data and significant manual effort to label objects. This limits their\napplicability in robotics, where solutions must scale to a large number of\nobjects and variety of conditions. This work proposes an autonomous process for\ntraining a Convolutional Neural Network (CNN) for object detection and pose\nestimation in robotic setups. The focus is on detecting objects placed in\ncluttered, tight environments, such as a shelf with multiple objects. In\nparticular, given access to 3D object models, several aspects of the\nenvironment are physically simulated. The models are placed in physically\nrealistic poses with respect to their environment to generate a labeled\nsynthetic dataset. To further improve object detection, the network self-trains\nover real images that are labeled using a robust multi-view pose estimation\nprocess. The proposed training process is evaluated on several existing\ndatasets and on a dataset collected for this paper with a Motoman robotic arm.\nResults show that the proposed approach outperforms popular training processes\nrelying on synthetic - but not physically realistic - data and manual\nannotation. The key contributions are the incorporation of physical reasoning\nin the synthetic data generation process and the automation of the annotation\nprocess over real images.","url_abs":"http://arxiv.org/abs/1703.03347v2","url_pdf":"http://arxiv.org/pdf/1703.03347v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"a-self-supervised-learning-system-for-object","repo_url":"https://github.com/cmitash/physim-dataset-generator","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"object","task_name":"Object"},{"task_slug":"object-detection","task_name":"Object Detection"},{"task_slug":"pose-estimation","task_name":"Pose Estimation"},{"task_slug":"self-supervised-learning","task_name":"Self-Supervised Learning"},{"task_slug":"synthetic-data-generation","task_name":"Synthetic Data Generation"},{"task_slug":"object-detection-1","task_name":"object-detection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1703.03347","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}