{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/incentivizing-exploration-in-reinforcement","title":"Incentivizing Exploration In Reinforcement Learning With Deep Predictive Models","arxiv_id":"1507.00814","date":"2015-07-03","proceeding":null,"authors":["Bradly C. Stadie","Sergey Levine","Pieter Abbeel"],"abstract":"Achieving efficient and scalable exploration in complex domains poses a major\nchallenge in reinforcement learning. While Bayesian and PAC-MDP approaches to\nthe exploration problem offer strong formal guarantees, they are often\nimpractical in higher dimensions due to their reliance on enumerating the\nstate-action space. Hence, exploration in complex domains is often performed\nwith simple epsilon-greedy methods. In this paper, we consider the challenging\nAtari games domain, which requires processing raw pixel inputs and delayed\nrewards. We evaluate several more sophisticated exploration strategies,\nincluding Thompson sampling and Boltzman exploration, and propose a new\nexploration method based on assigning exploration bonuses from a concurrently\nlearned model of the system dynamics. By parameterizing our learned model with\na neural network, we are able to develop a scalable and efficient approach to\nexploration bonuses that can be applied to tasks with complex, high-dimensional\nstate spaces. In the Atari domain, our method provides the most consistent\nimprovement across a range of games that pose a major challenge for prior\nmethods. In addition to raw game-scores, we also develop an AUC-100 metric for\nthe Atari Learning domain to evaluate the impact of exploration on this\nbenchmark.","url_abs":"http://arxiv.org/abs/1507.00814v3","url_pdf":"http://arxiv.org/pdf/1507.00814v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"incentivizing-exploration-in-reinforcement","repo_url":"https://github.com/CoffeeddCat/Multiagent_chainMDP","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}}],"tasks":[{"task_slug":"atari-games","task_name":"Atari Games"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"thompson-sampling","task_name":"Thompson Sampling"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/atari-games-on-atari-2600-freeway","task":"Atari Games","dataset":"Atari 2600 Freeway","model":"MP-EB","rank_in_archive_order":42,"of":59,"metrics":{"Score":"27.0"},"uses_additional_data":false},{"leaderboard":"/sota/atari-games-on-atari-2600-frostbite","task":"Atari Games","dataset":"Atari 2600 Frostbite","model":"MP-EB","rank_in_archive_order":38,"of":53,"metrics":{"Score":"507.0"},"uses_additional_data":false},{"leaderboard":"/sota/atari-games-on-atari-2600-montezumas-revenge","task":"Atari Games","dataset":"Atari 2600 Montezuma's Revenge","model":"MP-EB","rank_in_archive_order":25,"of":50,"metrics":{"Score":"142"},"uses_additional_data":false},{"leaderboard":"/sota/atari-games-on-atari-2600-qbert","task":"Atari Games","dataset":"Atari 2600 Q*Bert","model":"MP-EB","rank_in_archive_order":24,"of":57,"metrics":{"Score":"15805"},"uses_additional_data":false},{"leaderboard":"/sota/atari-games-on-atari-2600-venture","task":"Atari Games","dataset":"Atari 2600 Venture","model":"MP-EB","rank_in_archive_order":49,"of":55,"metrics":{"Score":"0.0"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1507.00814","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}