{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/deep-spatial-autoencoders-for-visuomotor","title":"Deep Spatial Autoencoders for Visuomotor Learning","arxiv_id":"1509.06113","date":"2015-09-21","proceeding":null,"authors":["Chelsea Finn","Xin Yu Tan","Yan Duan","Trevor Darrell","Sergey Levine","Pieter Abbeel"],"abstract":"Reinforcement learning provides a powerful and flexible framework for\nautomated acquisition of robotic motion skills. However, applying reinforcement\nlearning requires a sufficiently detailed representation of the state,\nincluding the configuration of task-relevant objects. We present an approach\nthat automates state-space construction by learning a state representation\ndirectly from camera images. Our method uses a deep spatial autoencoder to\nacquire a set of feature points that describe the environment for the current\ntask, such as the positions of objects, and then learns a motion skill with\nthese feature points using an efficient reinforcement learning method based on\nlocal linear models. The resulting controller reacts continuously to the\nlearned feature points, allowing the robot to dynamically manipulate objects in\nthe world with closed-loop control. We demonstrate our method with a PR2 robot\non tasks that include pushing a free-standing toy block, picking up a bag of\nrice using a spatula, and hanging a loop of rope on a hook at various\npositions. In each task, our method automatically learns to track task-relevant\nobjects and manipulate their configuration with the robot's arm.","url_abs":"http://arxiv.org/abs/1509.06113v3","url_pdf":"http://arxiv.org/pdf/1509.06113v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"deep-spatial-autoencoders-for-visuomotor","repo_url":"https://github.com/gorosgobe/dsae-torch","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1509.06113","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}