{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/variational-intrinsic-control","title":"Variational Intrinsic Control","arxiv_id":"1611.07507","date":"2016-11-22","proceeding":null,"authors":["Karol Gregor","Danilo Jimenez Rezende","Daan Wierstra"],"abstract":"In this paper we introduce a new unsupervised reinforcement learning method\nfor discovering the set of intrinsic options available to an agent. This set is\nlearned by maximizing the number of different states an agent can reliably\nreach, as measured by the mutual information between the set of options and\noption termination states. To this end, we instantiate two policy gradient\nbased algorithms, one that creates an explicit embedding space of options and\none that represents options implicitly. The algorithms also provide an explicit\nmeasure of empowerment in a given state that can be used by an empowerment\nmaximizing agent. The algorithm scales well with function approximation and we\ndemonstrate the applicability of the algorithm on a range of tasks.","url_abs":"http://arxiv.org/abs/1611.07507v1","url_pdf":"http://arxiv.org/pdf/1611.07507v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"variational-intrinsic-control","repo_url":"https://github.com/jbinas/gym-mnist","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"unsupervised-reinforcement-learning","task_name":"Unsupervised Reinforcement Learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1611.07507","atlas_url":"https://app.syntology.ai/?focus=1611.07507","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}