{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/information-directed-exploration-for-deep","title":"Information-Directed Exploration for Deep Reinforcement Learning","arxiv_id":"1812.07544","date":"2018-12-18","proceeding":"ICLR 2019 5","authors":["Nikolay Nikolov","Johannes Kirschner","Felix Berkenkamp","Andreas Krause"],"abstract":"Efficient exploration remains a major challenge for reinforcement learning.\nOne reason is that the variability of the returns often depends on the current\nstate and action, and is therefore heteroscedastic. Classical exploration\nstrategies such as upper confidence bound algorithms and Thompson sampling fail\nto appropriately account for heteroscedasticity, even in the bandit setting.\nMotivated by recent findings that address this issue in bandits, we propose to\nuse Information-Directed Sampling (IDS) for exploration in reinforcement\nlearning. As our main contribution, we build on recent advances in\ndistributional reinforcement learning and propose a novel, tractable\napproximation of IDS for deep Q-learning. The resulting exploration strategy\nexplicitly accounts for both parametric uncertainty and heteroscedastic\nobservation noise. We evaluate our method on Atari games and demonstrate a\nsignificant improvement over alternative approaches.","url_abs":"http://arxiv.org/abs/1812.07544v2","url_pdf":"http://arxiv.org/pdf/1812.07544v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"information-directed-exploration-for-deep","repo_url":"https://github.com/nikonikolov/rltf","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"atari-games","task_name":"Atari Games"},{"task_slug":"deep-reinforcement-learning","task_name":"Deep Reinforcement Learning"},{"task_slug":"distributional-reinforcement-learning","task_name":"Distributional Reinforcement Learning"},{"task_slug":"efficient-exploration","task_name":"Efficient Exploration"},{"task_slug":"q-learning","task_name":"Q-Learning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"thompson-sampling","task_name":"Thompson Sampling"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1812.07544","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}