{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/count-based-exploration-in-feature-space-for","title":"Count-Based Exploration in Feature Space for Reinforcement Learning","arxiv_id":"1706.08090","date":"2017-06-25","proceeding":null,"authors":["Jarryd Martin","Suraj Narayanan Sasikumar","Tom Everitt","Marcus Hutter"],"abstract":"We introduce a new count-based optimistic exploration algorithm for\nReinforcement Learning (RL) that is feasible in environments with\nhigh-dimensional state-action spaces. The success of RL algorithms in these\ndomains depends crucially on generalisation from limited training experience.\nFunction approximation techniques enable RL agents to generalise in order to\nestimate the value of unvisited states, but at present few methods enable\ngeneralisation regarding uncertainty. This has prevented the combination of\nscalable RL algorithms with efficient exploration strategies that drive the\nagent to reduce its uncertainty. We present a new method for computing a\ngeneralised state visit-count, which allows the agent to estimate the\nuncertainty associated with any state. Our \\phi-pseudocount achieves\ngeneralisation by exploiting same feature representation of the state space\nthat is used for value function approximation. States that have less frequently\nobserved features are deemed more uncertain. The \\phi-Exploration-Bonus\nalgorithm rewards the agent for exploring in feature space rather than in the\nuntransformed state space. The method is simpler and less computationally\nexpensive than some previous proposals, and achieves near state-of-the-art\nresults on high-dimensional RL benchmarks.","url_abs":"http://arxiv.org/abs/1706.08090v1","url_pdf":"http://arxiv.org/pdf/1706.08090v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"count-based-exploration-in-feature-space-for","repo_url":"https://github.com/aslanides/aixijs","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"atari-games","task_name":"Atari Games"},{"task_slug":"efficient-exploration","task_name":"Efficient Exploration"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/atari-games-on-atari-2600-freeway","task":"Atari Games","dataset":"Atari 2600 Freeway","model":"Sarsa-ε","rank_in_archive_order":34,"of":59,"metrics":{"Score":"29.9"},"uses_additional_data":false},{"leaderboard":"/sota/atari-games-on-atari-2600-freeway","task":"Atari Games","dataset":"Atari 2600 Freeway","model":"Sarsa-φ-EB","rank_in_archive_order":58,"of":59,"metrics":{"Score":"0.0"},"uses_additional_data":false},{"leaderboard":"/sota/atari-games-on-atari-2600-frostbite","task":"Atari Games","dataset":"Atari 2600 Frostbite","model":"Sarsa-φ-EB","rank_in_archive_order":27,"of":53,"metrics":{"Score":"2770.1"},"uses_additional_data":false},{"leaderboard":"/sota/atari-games-on-atari-2600-frostbite","task":"Atari Games","dataset":"Atari 2600 Frostbite","model":"Sarsa-ε","rank_in_archive_order":34,"of":53,"metrics":{"Score":"1394.3"},"uses_additional_data":false},{"leaderboard":"/sota/atari-games-on-atari-2600-montezumas-revenge","task":"Atari Games","dataset":"Atari 2600 Montezuma's Revenge","model":"Sarsa-φ-EB","rank_in_archive_order":13,"of":50,"metrics":{"Score":"2745.4"},"uses_additional_data":false},{"leaderboard":"/sota/atari-games-on-atari-2600-montezumas-revenge","task":"Atari Games","dataset":"Atari 2600 Montezuma's Revenge","model":"Sarsa-ε","rank_in_archive_order":22,"of":50,"metrics":{"Score":"399.5"},"uses_additional_data":false},{"leaderboard":"/sota/atari-games-on-atari-2600-qbert","task":"Atari Games","dataset":"Atari 2600 Q*Bert","model":"Sarsa-φ-EB","rank_in_archive_order":46,"of":57,"metrics":{"Score":"4111.8"},"uses_additional_data":false},{"leaderboard":"/sota/atari-games-on-atari-2600-qbert","task":"Atari Games","dataset":"Atari 2600 Q*Bert","model":"Sarsa-ε","rank_in_archive_order":47,"of":57,"metrics":{"Score":"3895.3"},"uses_additional_data":false},{"leaderboard":"/sota/atari-games-on-atari-2600-venture","task":"Atari Games","dataset":"Atari 2600 Venture","model":"Sarsa-φ-EB","rank_in_archive_order":18,"of":55,"metrics":{"Score":"1169.2"},"uses_additional_data":false},{"leaderboard":"/sota/atari-games-on-atari-2600-venture","task":"Atari Games","dataset":"Atari 2600 Venture","model":"Sarsa-ε","rank_in_archive_order":51,"of":55,"metrics":{"Score":"0.0"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1706.08090","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}