{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/multi-view-self-supervised-deep-learning-for","title":"Multi-view Self-supervised Deep Learning for 6D Pose Estimation in the Amazon Picking Challenge","arxiv_id":"1609.09475","date":"2016-09-29","proceeding":null,"authors":["Andy Zeng","Kuan-Ting Yu","Shuran Song","Daniel Suo","Ed Walker Jr.","Alberto Rodriguez","Jianxiong Xiao"],"abstract":"Robot warehouse automation has attracted significant interest in recent\nyears, perhaps most visibly in the Amazon Picking Challenge (APC). A fully\nautonomous warehouse pick-and-place system requires robust vision that reliably\nrecognizes and locates objects amid cluttered environments, self-occlusions,\nsensor noise, and a large variety of objects. In this paper we present an\napproach that leverages multi-view RGB-D data and self-supervised, data-driven\nlearning to overcome those difficulties. The approach was part of the\nMIT-Princeton Team system that took 3rd- and 4th- place in the stowing and\npicking tasks, respectively at APC 2016. In the proposed approach, we segment\nand label multiple views of a scene with a fully convolutional neural network,\nand then fit pre-scanned 3D object models to the resulting segmentation to get\nthe 6D object pose. Training a deep neural network for segmentation typically\nrequires a large amount of training data. We propose a self-supervised method\nto generate a large labeled dataset without tedious manual segmentation. We\ndemonstrate that our system can reliably estimate the 6D pose of objects under\na variety of scenarios. All code, data, and benchmarks are available at\nhttp://apc.cs.princeton.edu/","url_abs":"http://arxiv.org/abs/1609.09475v3","url_pdf":"http://arxiv.org/pdf/1609.09475v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"multi-view-self-supervised-deep-learning-for","repo_url":"https://github.com/andyzeng/apc-vision-toolbox","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"BSD-2-Clause"}},{"paper_slug":"multi-view-self-supervised-deep-learning-for","repo_url":"https://github.com/hz-ants/apc-vision-toolbox","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"BSD-2-Clause"}}],"tasks":[{"task_slug":"6d-pose-estimation-1","task_name":"6D Pose Estimation"},{"task_slug":"6d-pose-estimation-using-rgbd","task_name":"6D Pose Estimation using RGBD"},{"task_slug":"pose-estimation","task_name":"Pose Estimation"},{"task_slug":"segmentation","task_name":"Segmentation"}],"methods":[],"datasets_introduced":[{"slug":"shelf-tote","name":"Shelf&Tote Benchmark Dataset","full_name":"MIT-Princeton Amazon Picking Challenge 2016 Shelf&Tote Benchmark Dataset"},{"slug":"shelf-tote-training-dataset","name":"Shelf&Tote Training Dataset","full_name":"MIT-Princeton Amazon Picking Challenge 2016 Shelf&Tote Training Dataset"}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1609.09475","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}