{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/reset-free-trial-and-error-learning-for-robot","title":"Reset-free Trial-and-Error Learning for Robot Damage Recovery","arxiv_id":"1610.04213","date":"2016-10-13","proceeding":null,"authors":["Konstantinos Chatzilygeroudis","Vassilis Vassiliades","Jean-Baptiste Mouret"],"abstract":"The high probability of hardware failures prevents many advanced robots\n(e.g., legged robots) from being confidently deployed in real-world situations\n(e.g., post-disaster rescue). Instead of attempting to diagnose the failures,\nrobots could adapt by trial-and-error in order to be able to complete their\ntasks. In this situation, damage recovery can be seen as a Reinforcement\nLearning (RL) problem. However, the best RL algorithms for robotics require the\nrobot and the environment to be reset to an initial state after each episode,\nthat is, the robot is not learning autonomously. In addition, most of the RL\nmethods for robotics do not scale well with complex robots (e.g., walking\nrobots) and either cannot be used at all or take too long to converge to a\nsolution (e.g., hours of learning). In this paper, we introduce a novel\nlearning algorithm called \"Reset-free Trial-and-Error\" (RTE) that (1) breaks\nthe complexity by pre-generating hundreds of possible behaviors with a dynamics\nsimulator of the intact robot, and (2) allows complex robots to quickly recover\nfrom damage while completing their tasks and taking the environment into\naccount. We evaluate our algorithm on a simulated wheeled robot, a simulated\nsix-legged robot, and a real six-legged walking robot that are damaged in\nseveral ways (e.g., a missing leg, a shortened leg, faulty motor, etc.) and\nwhose objective is to reach a sequence of targets in an arena. Our experiments\nshow that the robots can recover most of their locomotion abilities in an\nenvironment with obstacles, and without any human intervention.","url_abs":"http://arxiv.org/abs/1610.04213v4","url_pdf":"http://arxiv.org/pdf/1610.04213v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"reset-free-trial-and-error-learning-for-robot","repo_url":"https://github.com/resibots/chatzilygeroudis_2018_rte","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}}],"tasks":[{"task_slug":"rte","task_name":"RTE"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1610.04213","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}