{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/leave-no-trace-learning-to-reset-for-safe-and","title":"Leave no Trace: Learning to Reset for Safe and Autonomous Reinforcement Learning","arxiv_id":"1711.06782","date":"2017-11-18","proceeding":"ICLR 2018 1","authors":["Benjamin Eysenbach","Shixiang Gu","Julian Ibarz","Sergey Levine"],"abstract":"Deep reinforcement learning algorithms can learn complex behavioral skills,\nbut real-world application of these methods requires a large amount of\nexperience to be collected by the agent. In practical settings, such as\nrobotics, this involves repeatedly attempting a task, resetting the environment\nbetween each attempt. However, not all tasks are easily or automatically\nreversible. In practice, this learning process requires extensive human\nintervention. In this work, we propose an autonomous method for safe and\nefficient reinforcement learning that simultaneously learns a forward and reset\npolicy, with the reset policy resetting the environment for a subsequent\nattempt. By learning a value function for the reset policy, we can\nautomatically determine when the forward policy is about to enter a\nnon-reversible state, providing for uncertainty-aware safety aborts. Our\nexperiments illustrate that proper use of the reset policy can greatly reduce\nthe number of manual resets required to learn a task, can reduce the number of\nunsafe actions that lead to non-reversible states, and can automatically induce\na curriculum.","url_abs":"http://arxiv.org/abs/1711.06782v1","url_pdf":"http://arxiv.org/pdf/1711.06782v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"leave-no-trace-learning-to-reset-for-safe-and","repo_url":"https://github.com/brain-research/LeaveNoTrace","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"deep-reinforcement-learning","task_name":"Deep Reinforcement Learning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1711.06782","atlas_url":"https://app.syntology.ai/?focus=1711.06782","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}