{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/contingency-aware-exploration-in","title":"Contingency-Aware Exploration in Reinforcement Learning","arxiv_id":"1811.01483","date":"2018-11-05","proceeding":"ICLR 2019 5","authors":["Jongwook Choi","Yijie Guo","Marcin Moczulski","Junhyuk Oh","Neal Wu","Mohammad Norouzi","Honglak Lee"],"abstract":"This paper investigates whether learning contingency-awareness and\ncontrollable aspects of an environment can lead to better exploration in\nreinforcement learning. To investigate this question, we consider an\ninstantiation of this hypothesis evaluated on the Arcade Learning Element\n(ALE). In this study, we develop an attentive dynamics model (ADM) that\ndiscovers controllable elements of the observations, which are often associated\nwith the location of the character in Atari games. The ADM is trained in a\nself-supervised fashion to predict the actions taken by the agent. The learned\ncontingency information is used as a part of the state representation for\nexploration purposes. We demonstrate that combining actor-critic algorithm with\ncount-based exploration using our representation achieves impressive results on\na set of notoriously challenging Atari games due to sparse rewards. For\nexample, we report a state-of-the-art score of >11,000 points on Montezuma's\nRevenge without using expert demonstrations, explicit high-level information\n(e.g., RAM states), or supervisory data. Our experiments confirm that\ncontingency-awareness is indeed an extremely powerful concept for tackling\nexploration problems in reinforcement learning and opens up interesting\nresearch questions for further investigations.","url_abs":"http://arxiv.org/abs/1811.01483v3","url_pdf":"http://arxiv.org/pdf/1811.01483v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"atari-games","task_name":"Atari Games"},{"task_slug":"montezumas-revenge","task_name":"Montezuma's Revenge"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/atari-games-on-atari-2600-montezumas-revenge","task":"Atari Games","dataset":"Atari 2600 Montezuma's Revenge","model":"A2C+CoEX","rank_in_archive_order":8,"of":50,"metrics":{"Score":"6635"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1811.01483","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}