{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/heuristics-answer-set-programming-and-markov","title":"Heuristics, Answer Set Programming and Markov Decision Process for Solving a Set of Spatial Puzzles","arxiv_id":"1903.03411","date":"2019-02-16","proceeding":null,"authors":["Thiago Freitas dos Santos","Paulo E. Santos","Leonardo A. Ferreira","Reinaldo A. C. Bianchi","Pedro Cabalar"],"abstract":"Spatial puzzles composed of rigid objects, flexible strings and holes offer\ninteresting domains for reasoning about spatial entities that are common in the\nhuman daily-life's activities. The goal of this work is to investigate the\nautomated solution of this kind of puzzles adapting an algorithm that combines\nAnswer Set Programming (ASP) with Markov Decision Process (MDP), algorithm\noASP(MDP), to use heuristics accelerating the learning process. ASP is applied\nto represent the domain as an MDP, while a Reinforcement Learning algorithm\n(Q-Learning) is used to find the optimal policies. In this work, the heuristics\nwere obtained from the solution of relaxed versions of the puzzles. Experiments\nwere performed on deterministic, non-deterministic and non-stationary versions\nof the puzzles. Results show that the proposed approach can accelerate the\nlearning process, presenting an advantage when compared to the non-heuristic\nversions of oASP(MDP) and Q-Learning.","url_abs":"http://arxiv.org/abs/1903.03411v1","url_pdf":"http://arxiv.org/pdf/1903.03411v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"heuristics-answer-set-programming-and-markov","repo_url":"https://bitbucket.org/thiagomestrado/journalarticle","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"q-learning","task_name":"Q-Learning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"}],"methods":[{"method_slug":"q-learning","method_name":"Q-Learning"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}