{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-meta-mdp-approach-to-exploration-for","title":"A Meta-MDP Approach to Exploration for Lifelong Reinforcement Learning","arxiv_id":"1902.00843","date":"2019-02-03","proceeding":"NeurIPS 2019 12","authors":["Francisco M. Garcia","Philip S. Thomas"],"abstract":"In this paper we consider the problem of how a reinforcement learning agent\nthat is tasked with solving a sequence of reinforcement learning problems (a\nsequence of Markov decision processes) can use knowledge acquired early in its\nlifetime to improve its ability to solve new problems. We argue that previous\nexperience with similar problems can provide an agent with information about\nhow it should explore when facing a new but related problem. We show that the\nsearch for an optimal exploration strategy can be formulated as a reinforcement\nlearning problem itself and demonstrate that such strategy can leverage\npatterns found in the structure of related problems. We conclude with\nexperiments that show the benefits of optimizing an exploration strategy using\nour proposed approach.","url_abs":"http://arxiv.org/abs/1902.00843v1","url_pdf":"http://arxiv.org/pdf/1902.00843v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"a-meta-mdp-approach-to-exploration-for","repo_url":"https://github.com/fmaxgarcia/Meta-MDP","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1902.00843","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}