{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-to-schedule-communication-in-multi","title":"Learning to Schedule Communication in Multi-agent Reinforcement Learning","arxiv_id":"1902.01554","date":"2019-02-05","proceeding":"ICLR 2019 5","authors":["Daewoo Kim","Sangwoo Moon","David Hostallero","Wan Ju Kang","Taeyoung Lee","Kyunghwan Son","Yung Yi"],"abstract":"Many real-world reinforcement learning tasks require multiple agents to make\nsequential decisions under the agents' interaction, where well-coordinated\nactions among the agents are crucial to achieve the target goal better at these\ntasks. One way to accelerate the coordination effect is to enable multiple\nagents to communicate with each other in a distributed manner and behave as a\ngroup. In this paper, we study a practical scenario when (i) the communication\nbandwidth is limited and (ii) the agents share the communication medium so that\nonly a restricted number of agents are able to simultaneously use the medium,\nas in the state-of-the-art wireless networking standards. This calls for a\ncertain form of communication scheduling. In that regard, we propose a\nmulti-agent deep reinforcement learning framework, called SchedNet, in which\nagents learn how to schedule themselves, how to encode the messages, and how to\nselect actions based on received messages. SchedNet is capable of deciding\nwhich agents should be entitled to broadcasting their (encoded) messages, by\nlearning the importance of each agent's partially observed information. We\nevaluate SchedNet against multiple baselines under two different applications,\nnamely, cooperative communication and navigation, and predator-prey. Our\nexperiments show a non-negligible performance gap between SchedNet and other\nmechanisms such as the ones without communication and with vanilla scheduling\nmethods, e.g., round robin, ranging from 32% to 43%.","url_abs":"http://arxiv.org/abs/1902.01554v1","url_pdf":"http://arxiv.org/pdf/1902.01554v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learning-to-schedule-communication-in-multi","repo_url":"https://github.com/rhoowd/sched_net","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"deep-reinforcement-learning","task_name":"Deep Reinforcement Learning"},{"task_slug":"multi-agent-reinforcement-learning","task_name":"Multi-agent Reinforcement Learning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"scheduling","task_name":"Scheduling"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1902.01554","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}