{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/the-surprising-effectiveness-of-mappo-in","title":"The Surprising Effectiveness of PPO in Cooperative, Multi-Agent Games","arxiv_id":"2103.01955","date":"2021-03-02","proceeding":null,"authors":["Chao Yu","Akash Velu","Eugene Vinitsky","Jiaxuan Gao","Yu Wang","Alexandre Bayen","Yi Wu"],"abstract":"Proximal Policy Optimization (PPO) is a ubiquitous on-policy reinforcement learning algorithm but is significantly less utilized than off-policy learning algorithms in multi-agent settings. This is often due to the belief that PPO is significantly less sample efficient than off-policy methods in multi-agent systems. In this work, we carefully study the performance of PPO in cooperative multi-agent settings. We show that PPO-based multi-agent algorithms achieve surprisingly strong performance in four popular multi-agent testbeds: the particle-world environments, the StarCraft multi-agent challenge, Google Research Football, and the Hanabi challenge, with minimal hyperparameter tuning and without any domain-specific algorithmic modifications or architectures. Importantly, compared to competitive off-policy methods, PPO often achieves competitive or superior results in both final returns and sample efficiency. Finally, through ablation studies, we analyze implementation and hyperparameter factors that are critical to PPO's empirical performance, and give concrete practical suggestions regarding these factors. Our results show that when using these practices, simple PPO-based methods can be a strong baseline in cooperative multi-agent reinforcement learning. Source code is released at \\url{https://github.com/marlbenchmark/on-policy}.","url_abs":"https://arxiv.org/abs/2103.01955v4","url_pdf":"https://arxiv.org/pdf/2103.01955v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"the-surprising-effectiveness-of-mappo-in","repo_url":"https://github.com/marlbenchmark/on-policy","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}},{"paper_slug":"the-surprising-effectiveness-of-mappo-in","repo_url":"https://github.com/16444take/aope-sim","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}},{"paper_slug":"the-surprising-effectiveness-of-mappo-in","repo_url":"https://github.com/ZifanWu/Coordinated-PPO","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"the-surprising-effectiveness-of-mappo-in","repo_url":"https://github.com/anonymous-iclr22/trust-region-in-multi-agent-reinforcement-learning","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"the-surprising-effectiveness-of-mappo-in","repo_url":"https://github.com/carolinewang01/dm2","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}},{"paper_slug":"the-surprising-effectiveness-of-mappo-in","repo_url":"https://github.com/chauncygu/multi-agent-constrained-policy-optimisation","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"the-surprising-effectiveness-of-mappo-in","repo_url":"https://github.com/cyanrain7/trpo-in-marl","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"the-surprising-effectiveness-of-mappo-in","repo_url":"https://github.com/cyanrain7/trust-region-policy-optimisation-in-multi-agent-reinforcement-learning","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"the-surprising-effectiveness-of-mappo-in","repo_url":"https://github.com/dig-beihang/ami","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"Apache-2.0"}},{"paper_slug":"the-surprising-effectiveness-of-mappo-in","repo_url":"https://github.com/emerge-lab/nocturne_lab","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"the-surprising-effectiveness-of-mappo-in","repo_url":"https://github.com/facebookresearch/benchmarl","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"the-surprising-effectiveness-of-mappo-in","repo_url":"https://github.com/ifpen/wfcrl-benchmark","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"the-surprising-effectiveness-of-mappo-in","repo_url":"https://github.com/morning9393/HAPPO-HATRPO","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"the-surprising-effectiveness-of-mappo-in","repo_url":"https://github.com/nju-rl/acorm","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"the-surprising-effectiveness-of-mappo-in","repo_url":"https://github.com/opendilab/DI-engine","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"the-surprising-effectiveness-of-mappo-in","repo_url":"https://github.com/tjuhaoxiaotian/pymarl3","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}},{"paper_slug":"the-surprising-effectiveness-of-mappo-in","repo_url":"https://github.com/tumcps/commonpower","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}},{"paper_slug":"the-surprising-effectiveness-of-mappo-in","repo_url":"https://github.com/zoeyuchao/mappo","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"the-surprising-effectiveness-of-mappo-in","repo_url":"https://github.com/xiuyu0000/new_papers_codes/tree/main/mappo","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null}],"tasks":[{"task_slug":"multi-agent-reinforcement-learning","task_name":"Multi-agent Reinforcement Learning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"starcraft","task_name":"Starcraft"},{"task_slug":"starcraft-ii","task_name":"Starcraft II"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[{"method_slug":"entropy-regularization","method_name":"Entropy Regularization"},{"method_slug":"ppo","method_name":"PPO"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2103.01955","atlas_url":"https://app.syntology.ai/?focus=2103.01955","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2103.01955"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"deterministic:regex_extraction","url":"https://github.com/marlbenchmark/on-policy","reach":{"status":"unanswered"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/zoeyuchao/mappo","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/tjuhaoxiaotian/pymarl3","reach":{"status":"unanswered"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/nju-rl/acorm","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/cyanrain7/trust-region-policy-optimisation-in-multi-agent-reinforcement-learning","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/dig-beihang/ami","reach":{"status":"ok","spdx":"Apache-2.0"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/cyanrain7/trpo-in-marl","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/16444take/aope-sim","reach":{"status":"unanswered"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/ifpen/wfcrl-benchmark","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/carolinewang01/dm2","reach":{"status":"unanswered"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/chauncygu/multi-agent-constrained-policy-optimisation","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/anonymous-iclr22/trust-region-in-multi-agent-reinforcement-learning","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/tumcps/commonpower","reach":{"status":"unanswered"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/ZifanWu/Coordinated-PPO","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/xiuyu0000/new_papers_codes/tree/main/mappo","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/opendilab/DI-engine","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/morning9393/HAPPO-HATRPO","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/emerge-lab/nocturne_lab","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/facebookresearch/benchmarl","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"ran_draft_wrong":2},"by_repo_kind":{"listed":{"samples":2,"ran":2,"repositories":2}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"4cee71b4d9b5a737","entry":"linear_schedule","repo":"emerge-lab/nocturne_lab","repo_kind":"listed","path":"experiments/hr_rl/run_hr_ppo_cli.py","file_url":"https://github.com/emerge-lab/nocturne_lab/blob/HEAD/experiments/hr_rl/run_hr_ppo_cli.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"4cee71b4d9b5a737"}},{"code_sha256_prefix":"3d51ddb42232f4be","entry":"parse_args","repo":"ZifanWu/Coordinated-PPO","repo_kind":"listed","path":"onpolicy/scripts/train/train_smac.py","file_url":"https://github.com/ZifanWu/Coordinated-PPO/blob/HEAD/onpolicy/scripts/train/train_smac.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"3d51ddb42232f4be"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}