{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/logically-constrained-reinforcement-learning","title":"Logically-Constrained Reinforcement Learning","arxiv_id":"1801.08099","date":"2018-01-24","proceeding":null,"authors":["Mohammadhosein Hasanbeig","Alessandro Abate","Daniel Kroening"],"abstract":"We present the first model-free Reinforcement Learning (RL) algorithm to\nsynthesise policies for an unknown Markov Decision Process (MDP), such that a\nlinear time property is satisfied. The given temporal property is converted\ninto a Limit Deterministic Buchi Automaton (LDBA) and a robust reward function\nis defined over the state-action pairs of the MDP according to the resulting\nLDBA. With this reward function, the policy synthesis procedure is\n\"constrained\" by the given specification. These constraints guide the MDP\nexploration so as to minimize the solution time by only considering the portion\nof the MDP that is relevant to satisfaction of the LTL property. This improves\nperformance and scalability of the proposed method by avoiding an exhaustive\nupdate over the whole state space while the efficiency of standard methods such\nas dynamic programming is hindered by excessive memory requirements, caused by\nthe need to store a full-model in memory. Additionally, we show that the RL\nprocedure sets up a local value iteration method to efficiently calculate the\nmaximum probability of satisfying the given property, at any given state of the\nMDP. We prove that our algorithm is guaranteed to find a policy whose traces\nprobabilistically satisfy the LTL property if such a policy exists, and\nadditionally we show that our method produces reasonable control policies even\nwhen the LTL property cannot be satisfied. The performance of the algorithm is\nevaluated via a set of numerical examples. We observe an improvement of one\norder of magnitude in the number of iterations required for the synthesis\ncompared to existing approaches.","url_abs":"http://arxiv.org/abs/1801.08099v8","url_pdf":"http://arxiv.org/pdf/1801.08099v8.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"logically-constrained-reinforcement-learning","repo_url":"https://github.com/grockious/lcrl","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"decision-making","task_name":"Decision Making"},{"task_slug":"decision-making-under-uncertainty","task_name":"Decision Making Under Uncertainty"},{"task_slug":"hierarchical-reinforcement-learning","task_name":"Hierarchical Reinforcement Learning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"safe-reinforcement-learning","task_name":"Safe Reinforcement Learning"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1801.08099","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}