{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/end-to-end-safe-reinforcement-learning","title":"End-to-End Safe Reinforcement Learning through Barrier Functions for Safety-Critical Continuous Control Tasks","arxiv_id":"1903.08792","date":"2019-03-21","proceeding":null,"authors":["Richard Cheng","Gabor Orosz","Richard M. Murray","Joel W. Burdick"],"abstract":"Reinforcement Learning (RL) algorithms have found limited success beyond\nsimulated applications, and one main reason is the absence of safety guarantees\nduring the learning process. Real world systems would realistically fail or\nbreak before an optimal controller can be learned. To address this issue, we\npropose a controller architecture that combines (1) a model-free RL-based\ncontroller with (2) model-based controllers utilizing control barrier functions\n(CBFs) and (3) on-line learning of the unknown system dynamics, in order to\nensure safety during learning. Our general framework leverages the success of\nRL algorithms to learn high-performance controllers, while the CBF-based\ncontrollers both guarantee safety and guide the learning process by\nconstraining the set of explorable polices. We utilize Gaussian Processes (GPs)\nto model the system dynamics and its uncertainties.\n  Our novel controller synthesis algorithm, RL-CBF, guarantees safety with high\nprobability during the learning process, regardless of the RL algorithm used,\nand demonstrates greater policy exploration efficiency. We test our algorithm\non (1) control of an inverted pendulum and (2) autonomous car-following with\nwireless vehicle-to-vehicle communication, and show that our algorithm attains\nmuch greater sample efficiency in learning than other state-of-the-art\nalgorithms and maintains safety during the entire learning process.","url_abs":"http://arxiv.org/abs/1903.08792v1","url_pdf":"http://arxiv.org/pdf/1903.08792v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"end-to-end-safe-reinforcement-learning","repo_url":"https://github.com/rcheng805/RL-CBF","is_official":0,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"tf","reach":{"status":"ok"}}],"tasks":[{"task_slug":"continuous-control","task_name":"Continuous Control"},{"task_slug":"gaussian-processes","task_name":"Gaussian Processes"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"safe-reinforcement-learning","task_name":"Safe Reinforcement Learning"},{"task_slug":"continuous-control","task_name":"continuous-control"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1903.08792","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}