{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/safe-reinforcement-learning-via-shielding","title":"Safe Reinforcement Learning via Shielding","arxiv_id":"1708.08611","date":"2017-08-29","proceeding":null,"authors":["Mohammed Alshiekh","Roderick Bloem","Ruediger Ehlers","Bettina Könighofer","Scott Niekum","Ufuk Topcu"],"abstract":"Reinforcement learning algorithms discover policies that maximize reward, but\ndo not necessarily guarantee safety during learning or execution phases. We\nintroduce a new approach to learn optimal policies while enforcing properties\nexpressed in temporal logic. To this end, given the temporal logic\nspecification that is to be obeyed by the learning system, we propose to\nsynthesize a reactive system called a shield. The shield is introduced in the\ntraditional learning process in two alternative ways, depending on the location\nat which the shield is implemented. In the first one, the shield acts each time\nthe learning agent is about to make a decision and provides a list of safe\nactions. In the second way, the shield is introduced after the learning agent.\nThe shield monitors the actions from the learner and corrects them only if the\nchosen action causes a violation of the specification. We discuss which\nrequirements a shield must meet to preserve the convergence guarantees of the\nlearner. Finally, we demonstrate the versatility of our approach on several\nchallenging reinforcement learning scenarios.","url_abs":"http://arxiv.org/abs/1708.08611v2","url_pdf":"http://arxiv.org/pdf/1708.08611v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"safe-reinforcement-learning-via-shielding","repo_url":"https://github.com/DanielLSM/safe-rl-tutorial","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"safe-reinforcement-learning","task_name":"Safe Reinforcement Learning"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1708.08611","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}