{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/scalable-coordinated-exploration-in","title":"Scalable Coordinated Exploration in Concurrent Reinforcement Learning","arxiv_id":"1805.08948","date":"2018-05-23","proceeding":"NeurIPS 2018 12","authors":["Maria Dimakopoulou","Ian Osband","Benjamin Van Roy"],"abstract":"We consider a team of reinforcement learning agents that concurrently operate\nin a common environment, and we develop an approach to efficient coordinated\nexploration that is suitable for problems of practical scale. Our approach\nbuilds on seed sampling (Dimakopoulou and Van Roy, 2018) and randomized value\nfunction learning (Osband et al., 2016). We demonstrate that, for simple\ntabular contexts, the approach is competitive with previously proposed tabular\nmodel learning methods (Dimakopoulou and Van Roy, 2018). With a\nhigher-dimensional problem and a neural network value function representation,\nthe approach learns quickly with far fewer agents than alternative exploration\nschemes.","url_abs":"http://arxiv.org/abs/1805.08948v2","url_pdf":"http://arxiv.org/pdf/1805.08948v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"scalable-coordinated-exploration-in","repo_url":"https://github.com/efancher/cs234_work_final_project","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1805.08948","atlas_url":"https://app.syntology.ai/?focus=1805.08948","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}