{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/universal-successor-features-approximators","title":"Universal Successor Features Approximators","arxiv_id":"1812.07626","date":"2018-12-18","proceeding":"ICLR 2019 5","authors":["Diana Borsa","André Barreto","John Quan","Daniel Mankowitz","Rémi Munos","Hado van Hasselt","David Silver","Tom Schaul"],"abstract":"The ability of a reinforcement learning (RL) agent to learn about many reward\nfunctions at the same time has many potential benefits, such as the\ndecomposition of complex tasks into simpler ones, the exchange of information\nbetween tasks, and the reuse of skills. We focus on one aspect in particular,\nnamely the ability to generalise to unseen tasks. Parametric generalisation\nrelies on the interpolation power of a function approximator that is given the\ntask description as input; one of its most common form are universal value\nfunction approximators (UVFAs). Another way to generalise to new tasks is to\nexploit structure in the RL problem itself. Generalised policy improvement\n(GPI) combines solutions of previous tasks into a policy for the unseen task;\nthis relies on instantaneous policy evaluation of old policies under the new\nreward function, which is made possible through successor features (SFs). Our\nproposed universal successor features approximators (USFAs) combine the\nadvantages of all of these, namely the scalability of UVFAs, the instant\ninference of SFs, and the strong generalisation of GPI. We discuss the\nchallenges involved in training a USFA, its generalisation properties and\ndemonstrate its practical benefits and transfer abilities on a large-scale\ndomain in which the agent has to navigate in a first-person perspective\nthree-dimensional environment.","url_abs":"http://arxiv.org/abs/1812.07626v1","url_pdf":"http://arxiv.org/pdf/1812.07626v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"universal-successor-features-approximators","repo_url":"https://github.com/Wanqianxn/usfa","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"universal-successor-features-approximators","repo_url":"https://github.com/wcarvalho/jaxneurorl","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"jax","reach":{"status":"ok"}}],"tasks":[{"task_slug":"navigate","task_name":"Navigate"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1812.07626","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}