{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/lyapunov-based-safe-policy-optimization-for","title":"Lyapunov-based Safe Policy Optimization for Continuous Control","arxiv_id":"1901.10031","date":"2019-01-28","proceeding":null,"authors":["Yin-Lam Chow","Ofir Nachum","Aleksandra Faust","Edgar Duenez-Guzman","Mohammad Ghavamzadeh"],"abstract":"We study continuous action reinforcement learning problems in which it is\ncrucial that the agent interacts with the environment only through safe\npolicies, i.e.,~policies that do not take the agent to undesirable situations.\nWe formulate these problems as constrained Markov decision processes (CMDPs)\nand present safe policy optimization algorithms that are based on a Lyapunov\napproach to solve them. Our algorithms can use any standard policy gradient\n(PG) method, such as deep deterministic policy gradient (DDPG) or proximal\npolicy optimization (PPO), to train a neural network policy, while guaranteeing\nnear-constraint satisfaction for every policy update by projecting either the\npolicy parameter or the action onto the set of feasible solutions induced by\nthe state-dependent linearized Lyapunov constraints. Compared to the existing\nconstrained PG algorithms, ours are more data efficient as they are able to\nutilize both on-policy and off-policy data. Moreover, our action-projection\nalgorithm often leads to less conservative policy updates and allows for\nnatural integration into an end-to-end PG training pipeline. We evaluate our\nalgorithms and compare them with the state-of-the-art baselines on several\nsimulated (MuJoCo) tasks, as well as a real-world indoor robot navigation\nproblem, demonstrating their effectiveness in terms of balancing performance\nand constraint satisfaction. Videos of the experiments can be found in the\nfollowing link:\nhttps://drive.google.com/file/d/1pzuzFqWIE710bE2U6DmS59AfRzqK2Kek/view?usp=sharing.","url_abs":"http://arxiv.org/abs/1901.10031v2","url_pdf":"http://arxiv.org/pdf/1901.10031v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"lyapunov-based-safe-policy-optimization-for","repo_url":"https://github.com/jemaw/gym-safety","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"continuous-control","task_name":"Continuous Control"},{"task_slug":"mujoco","task_name":"MuJoCo"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"robot-navigation","task_name":"Robot Navigation"},{"task_slug":"continuous-control","task_name":"continuous-control"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1901.10031","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}