{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/feature-control-as-intrinsic-motivation-for","title":"Feature Control as Intrinsic Motivation for Hierarchical Reinforcement Learning","arxiv_id":"1705.06769","date":"2017-05-18","proceeding":null,"authors":["Nat Dilokthanakul","Christos Kaplanis","Nick Pawlowski","Murray Shanahan"],"abstract":"The problem of sparse rewards is one of the hardest challenges in\ncontemporary reinforcement learning. Hierarchical reinforcement learning (HRL)\ntackles this problem by using a set of temporally-extended actions, or options,\neach of which has its own subgoal. These subgoals are normally handcrafted for\nspecific tasks. Here, though, we introduce a generic class of subgoals with\nbroad applicability in the visual domain. Underlying our approach (in common\nwith work using \"auxiliary tasks\") is the hypothesis that the ability to\ncontrol aspects of the environment is an inherently useful skill to have. We\nincorporate such subgoals in an end-to-end hierarchical reinforcement learning\nsystem and test two variants of our algorithm on a number of games from the\nAtari suite. We highlight the advantage of our approach in one of the hardest\ngames -- Montezuma's revenge -- for which the ability to handle sparse rewards\nis key. Our agent learns several times faster than the current state-of-the-art\nHRL agent in this game, reaching a similar level of performance. UPDATE\n22/11/17: We found that a standard A3C agent with a simple shaped reward, i.e.\nextrinsic reward + feature control intrinsic reward, has comparable performance\nto our agent in Montezuma Revenge. In light of the new experiments performed,\nthe advantage of our HRL approach can be attributed more to its ability to\nlearn useful features from intrinsic rewards rather than its ability to explore\nand reuse abstracted skills with hierarchical components. This has led us to a\nnew conclusion about the result.","url_abs":"http://arxiv.org/abs/1705.06769v2","url_pdf":"http://arxiv.org/pdf/1705.06769v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"feature-control-as-intrinsic-motivation-for","repo_url":"https://github.com/lyebi/Test","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"hierarchical-reinforcement-learning","task_name":"Hierarchical Reinforcement Learning"},{"task_slug":"montezumas-revenge","task_name":"Montezuma's Revenge"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[{"method_slug":"a3c","method_name":"A3C"},{"method_slug":"convolution","method_name":"Convolution"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"entropy-regularization","method_name":"Entropy Regularization"},{"method_slug":"softmax","method_name":"Softmax"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1705.06769","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}