{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/deconfounding-reinforcement-learning-in","title":"Deconfounding Reinforcement Learning in Observational Settings","arxiv_id":"1812.10576","date":"2018-12-26","proceeding":null,"authors":["Chaochao Lu","Bernhard Schölkopf","José Miguel Hernández-Lobato"],"abstract":"We propose a general formulation for addressing reinforcement learning (RL)\nproblems in settings with observational data. That is, we consider the problem\nof learning good policies solely from historical data in which unobserved\nfactors (confounders) affect both observed actions and rewards. Our formulation\nallows us to extend a representative RL algorithm, the Actor-Critic method, to\nits deconfounding variant, with the methodology for this extension being easily\napplied to other RL algorithms. In addition to this, we develop a new benchmark\nfor evaluating deconfounding RL algorithms by modifying the OpenAI Gym\nenvironments and the MNIST dataset. Using this benchmark, we demonstrate that\nthe proposed algorithms are superior to traditional RL methods in confounded\nenvironments with observational data. To the best of our knowledge, this is the\nfirst time that confounders are taken into consideration for addressing full RL\nproblems with observational data. Code is available at\nhttps://github.com/CausalRL/DRL.","url_abs":"http://arxiv.org/abs/1812.10576v1","url_pdf":"http://arxiv.org/pdf/1812.10576v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"deconfounding-reinforcement-learning-in","repo_url":"https://github.com/CausalRL/DRL","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}}],"tasks":[{"task_slug":"openai-gym","task_name":"OpenAI Gym"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1812.10576","atlas_url":"https://app.syntology.ai/?focus=1812.10576","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}