{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/exit-oos-towards-learning-from-planning-in","title":"ExIt-OOS: Towards Learning from Planning in Imperfect Information Games","arxiv_id":"1808.10120","date":"2018-08-30","proceeding":null,"authors":["Andy Kitchen","Michela Benedetti"],"abstract":"The current state of the art in playing many important perfect information\ngames, including Chess and Go, combines planning and deep reinforcement\nlearning with self-play. We extend this approach to imperfect information games\nand present ExIt-OOS, a novel approach to playing imperfect information games\nwithin the Expert Iteration framework and inspired by AlphaZero. We use Online\nOutcome Sampling, an online search algorithm for imperfect information games in\nplace of MCTS. While training online, our neural strategy is used to improve\nthe accuracy of playouts in OOS, allowing a learning and planning feedback loop\nfor imperfect information games.","url_abs":"http://arxiv.org/abs/1808.10120v2","url_pdf":"http://arxiv.org/pdf/1808.10120v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"exit-oos-towards-learning-from-planning-in","repo_url":"https://github.com/IAARhub/TrucoAnalytics","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"deep-reinforcement-learning","task_name":"Deep Reinforcement Learning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[{"method_slug":"alphazero","method_name":"AlphaZero"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}