{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/concurrent-meta-reinforcement-learning","title":"Concurrent Meta Reinforcement Learning","arxiv_id":"1903.02710","date":"2019-03-07","proceeding":null,"authors":["Emilio Parisotto","Soham Ghosh","Sai Bhargav Yalamanchi","Varsha Chinnaobireddy","Yuhuai Wu","Ruslan Salakhutdinov"],"abstract":"State-of-the-art meta reinforcement learning algorithms typically assume the\nsetting of a single agent interacting with its environment in a sequential\nmanner. A negative side-effect of this sequential execution paradigm is that,\nas the environment becomes more and more challenging, and thus requiring more\ninteraction episodes for the meta-learner, it needs the agent to reason over\nlonger and longer time-scales. To combat the difficulty of long time-scale\ncredit assignment, we propose an alternative parallel framework, which we name\n\"Concurrent Meta-Reinforcement Learning\" (CMRL), that transforms the temporal\ncredit assignment problem into a multi-agent reinforcement learning one. In\nthis multi-agent setting, a set of parallel agents are executed in the same\nenvironment and each of these \"rollout\" agents are given the means to\ncommunicate with each other. The goal of the communication is to coordinate, in\na collaborative manner, the most efficient exploration of the shared task the\nagents are currently assigned. This coordination therefore represents the\nmeta-learning aspect of the framework, as each agent can be assigned or assign\nitself a particular section of the current task's state space. This framework\nis in contrast to standard RL methods that assume that each parallel rollout\noccurs independently, which can potentially waste computation if many of the\nrollouts end up sampling the same part of the state space. Furthermore, the\nparallel setting enables us to define several reward sharing functions and\nauxiliary losses that are non-trivial to apply in the sequential setting. We\ndemonstrate the effectiveness of our proposed CMRL at improving over sequential\nmethods in a variety of challenging tasks.","url_abs":"http://arxiv.org/abs/1903.02710v1","url_pdf":"http://arxiv.org/pdf/1903.02710v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"concurrent-meta-reinforcement-learning","repo_url":"https://github.com/impredicative/irc-rss-feed-bot","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"AGPL-3.0"}}],"tasks":[{"task_slug":"efficient-exploration","task_name":"Efficient Exploration"},{"task_slug":"meta-reinforcement-learning","task_name":"Meta Reinforcement Learning"},{"task_slug":"meta-learning","task_name":"Meta-Learning"},{"task_slug":"multi-agent-reinforcement-learning","task_name":"Multi-agent Reinforcement Learning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1903.02710","atlas_url":"https://app.syntology.ai/?focus=1903.02710","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}