{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/trust-region-policy-optimisation-in-multi","title":"Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning","arxiv_id":"2109.11251","date":"2021-09-23","proceeding":"ICLR 2022 4","authors":["Jakub Grudzien Kuba","Ruiqing Chen","Muning Wen","Ying Wen","Fanglei Sun","Jun Wang","Yaodong Yang"],"abstract":"Trust region methods rigorously enabled reinforcement learning (RL) agents to learn monotonically improving policies, leading to superior performance on a variety of tasks. Unfortunately, when it comes to multi-agent reinforcement learning (MARL), the property of monotonic improvement may not simply apply; this is because agents, even in cooperative games, could have conflicting directions of policy updates. As a result, achieving a guaranteed improvement on the joint policy where each agent acts individually remains an open challenge. In this paper, we extend the theory of trust region learning to MARL. Central to our findings are the multi-agent advantage decomposition lemma and the sequential policy update scheme. Based on these, we develop Heterogeneous-Agent Trust Region Policy Optimisation (HATPRO) and Heterogeneous-Agent Proximal Policy Optimisation (HAPPO) algorithms. Unlike many existing MARL algorithms, HATRPO/HAPPO do not need agents to share parameters, nor do they need any restrictive assumptions on decomposibility of the joint value function. Most importantly, we justify in theory the monotonic improvement property of HATRPO/HAPPO. We evaluate the proposed methods on a series of Multi-Agent MuJoCo and StarCraftII tasks. Results show that HATRPO and HAPPO significantly outperform strong baselines such as IPPO, MAPPO and MADDPG on all tested tasks, therefore establishing a new state of the art.","url_abs":"https://arxiv.org/abs/2109.11251v2","url_pdf":"https://arxiv.org/pdf/2109.11251v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"trust-region-policy-optimisation-in-multi","repo_url":"https://github.com/cyanrain7/trust-region-policy-optimisation-in-multi-agent-reinforcement-learning","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"trust-region-policy-optimisation-in-multi","repo_url":"https://github.com/anonymous-iclr22/trust-region-in-multi-agent-reinforcement-learning","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":null},{"paper_slug":"trust-region-policy-optimisation-in-multi","repo_url":"https://github.com/chauncygu/multi-agent-constrained-policy-optimisation","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"trust-region-policy-optimisation-in-multi","repo_url":"https://github.com/chauncygu/safe-multi-agent-mujoco","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"trust-region-policy-optimisation-in-multi","repo_url":"https://github.com/chauncygu/safe-multi-agent-robosuite","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"trust-region-policy-optimisation-in-multi","repo_url":"https://github.com/cyanrain7/trpo-in-marl","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"trust-region-policy-optimisation-in-multi","repo_url":"https://github.com/eduardosebastianrodriguez/phmarl","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"trust-region-policy-optimisation-in-multi","repo_url":"https://github.com/mehdinasiri/mirror-descent-in-marl","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"trust-region-policy-optimisation-in-multi","repo_url":"https://github.com/morning9393/HAPPO-HATRPO","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"trust-region-policy-optimisation-in-multi","repo_url":"https://github.com/pku-marl/multi-agent-transformer","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"trust-region-policy-optimisation-in-multi","repo_url":"https://github.com/opendilab/DI-engine","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"lemma","task_name":"LEMMA"},{"task_slug":"mujoco","task_name":"MuJoCo"},{"task_slug":"multi-agent-reinforcement-learning","task_name":"Multi-agent Reinforcement Learning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[{"method_slug":"adam","method_name":"Adam"},{"method_slug":"batch-normalization","method_name":"Batch Normalization"},{"method_slug":"convolution","method_name":"Convolution"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"experience-replay","method_name":"Experience Replay"},{"method_slug":"maddpg","method_name":"MADDPG"},{"method_slug":"relu","method_name":"ReLU"},{"method_slug":"weight-decay","method_name":"Weight Decay"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2109.11251","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2109.11251"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/pku-marl/multi-agent-transformer","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/eduardosebastianrodriguez/phmarl","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/cyanrain7/trust-region-policy-optimisation-in-multi-agent-reinforcement-learning","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/chauncygu/safe-multi-agent-mujoco","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/cyanrain7/trpo-in-marl","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/chauncygu/safe-multi-agent-robosuite","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/chauncygu/multi-agent-constrained-policy-optimisation","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/anonymous-iclr22/trust-region-in-multi-agent-reinforcement-learning","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/mehdinasiri/mirror-descent-in-marl","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/opendilab/DI-engine","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/morning9393/HAPPO-HATRPO","reach":null}],"summary":{"ran_honours":1,"ran_draft_wrong":1,"unverified":3},"by_repo_kind":{"listed":{"samples":5,"ran":2,"repositories":2}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"9589eca5bae9de80","entry":"check","repo":"mehdinasiri/mirror-descent-in-marl","repo_kind":"listed","path":"algorithms/utils/util.py","file_url":"https://github.com/mehdinasiri/mirror-descent-in-marl/blob/HEAD/algorithms/utils/util.py","link_basis":"harvester_set","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"9589eca5bae9de80"}},{"code_sha256_prefix":"9de355e93051e4ad","entry":"init","repo":"mehdinasiri/mirror-descent-in-marl","repo_kind":"listed","path":"algorithms/utils/util.py","file_url":"https://github.com/mehdinasiri/mirror-descent-in-marl/blob/HEAD/algorithms/utils/util.py","link_basis":"harvester_set","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"9de355e93051e4ad"}},{"code_sha256_prefix":"0944240e80ec7625","entry":"get_clones","repo":"mehdinasiri/mirror-descent-in-marl","repo_kind":"listed","path":"algorithms/utils/util.py","file_url":"https://github.com/mehdinasiri/mirror-descent-in-marl/blob/HEAD/algorithms/utils/util.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"0944240e80ec7625"}},{"code_sha256_prefix":"2b9d28f20991d3ec","entry":"laplacian","repo":"eduardosebastianrodriguez/phmarl","repo_kind":"listed","path":"robotarium/functions.py","file_url":"https://github.com/eduardosebastianrodriguez/phmarl/blob/HEAD/robotarium/functions.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"2b9d28f20991d3ec"}},{"code_sha256_prefix":"2fc6fdb85f13dc41","entry":"none_or_str","repo":"eduardosebastianrodriguez/phmarl","repo_kind":"listed","path":"parse_args.py","file_url":"https://github.com/eduardosebastianrodriguez/phmarl/blob/HEAD/parse_args.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"2fc6fdb85f13dc41"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}