{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-lyapunov-based-approach-to-safe","title":"A Lyapunov-based Approach to Safe Reinforcement Learning","arxiv_id":"1805.07708","date":"2018-05-20","proceeding":"NeurIPS 2018 12","authors":["Yin-Lam Chow","Ofir Nachum","Edgar Duenez-Guzman","Mohammad Ghavamzadeh"],"abstract":"In many real-world reinforcement learning (RL) problems, besides optimizing\nthe main objective function, an agent must concurrently avoid violating a\nnumber of constraints. In particular, besides optimizing performance it is\ncrucial to guarantee the safety of an agent during training as well as\ndeployment (e.g. a robot should avoid taking actions - exploratory or not -\nwhich irrevocably harm its hardware). To incorporate safety in RL, we derive\nalgorithms under the framework of constrained Markov decision problems (CMDPs),\nan extension of the standard Markov decision problems (MDPs) augmented with\nconstraints on expected cumulative costs. Our approach hinges on a novel\n\\emph{Lyapunov} method. We define and present a method for constructing\nLyapunov functions, which provide an effective way to guarantee the global\nsafety of a behavior policy during training via a set of local, linear\nconstraints. Leveraging these theoretical underpinnings, we show how to use the\nLyapunov approach to systematically transform dynamic programming (DP) and RL\nalgorithms into their safe counterparts. To illustrate their effectiveness, we\nevaluate these algorithms in several CMDP planning and decision-making tasks on\na safety benchmark domain. Our results show that our proposed method\nsignificantly outperforms existing baselines in balancing constraint\nsatisfaction and performance.","url_abs":"http://arxiv.org/abs/1805.07708v1","url_pdf":"http://arxiv.org/pdf/1805.07708v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"a-lyapunov-based-approach-to-safe","repo_url":"https://github.com/jemaw/gym-safety","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"decision-making","task_name":"Decision Making"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"safe-reinforcement-learning","task_name":"Safe Reinforcement Learning"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1805.07708","atlas_url":"https://app.syntology.ai/?focus=1805.07708","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}