{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/transferring-end-to-end-visuomotor-control","title":"Transferring End-to-End Visuomotor Control from Simulation to Real World for a Multi-Stage Task","arxiv_id":"1707.02267","date":"2017-07-07","proceeding":null,"authors":["Stephen James","Andrew J. Davison","Edward Johns"],"abstract":"End-to-end control for robot manipulation and grasping is emerging as an\nattractive alternative to traditional pipelined approaches. However, end-to-end\nmethods tend to either be slow to train, exhibit little or no generalisability,\nor lack the ability to accomplish long-horizon or multi-stage tasks. In this\npaper, we show how two simple techniques can lead to end-to-end (image to\nvelocity) execution of a multi-stage task, which is analogous to a simple\ntidying routine, without having seen a single real image. This involves\nlocating, reaching for, and grasping a cube, then locating a basket and\ndropping the cube inside. To achieve this, robot trajectories are computed in a\nsimulator, to collect a series of control velocities which accomplish the task.\nThen, a CNN is trained to map observed images to velocities, using domain\nrandomisation to enable generalisation to real world images. Results show that\nwe are able to successfully accomplish the task in the real world with the\nability to generalise to novel environments, including those with dynamic\nlighting conditions, distractor objects, and moving objects, including the\nbasket itself. We believe our approach to be simple, highly scalable, and\ncapable of learning long-horizon tasks that have until now not been shown with\nthe state-of-the-art in end-to-end robot control.","url_abs":"http://arxiv.org/abs/1707.02267v2","url_pdf":"http://arxiv.org/pdf/1707.02267v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"transferring-end-to-end-visuomotor-control","repo_url":"https://github.com/stepjam/PyRep","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"robot-manipulation","task_name":"Robot Manipulation"},{"task_slug":"robotic-grasping","task_name":"Robotic Grasping"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1707.02267","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}