{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/multi-agent-constrained-policy-optimisation","title":"Multi-Agent Constrained Policy Optimisation","arxiv_id":"2110.02793","date":"2021-10-06","proceeding":null,"authors":["Shangding Gu","Jakub Grudzien Kuba","Munning Wen","Ruiqing Chen","Ziyan Wang","Zheng Tian","Jun Wang","Alois Knoll","Yaodong Yang"],"abstract":"Developing reinforcement learning algorithms that satisfy safety constraints is becoming increasingly important in real-world applications. In multi-agent reinforcement learning (MARL) settings, policy optimisation with safety awareness is particularly challenging because each individual agent has to not only meet its own safety constraints, but also consider those of others so that their joint behaviour can be guaranteed safe. Despite its importance, the problem of safe multi-agent learning has not been rigorously studied; very few solutions have been proposed, nor a sharable testing environment or benchmarks. To fill these gaps, in this work, we formulate the safe MARL problem as a constrained Markov game and solve it with policy optimisation methods. Our solutions -- Multi-Agent Constrained Policy Optimisation (MACPO) and MAPPO-Lagrangian -- leverage the theories from both constrained policy optimisation and multi-agent trust region learning. Crucially, our methods enjoy theoretical guarantees of both monotonic improvement in reward and satisfaction of safety constraints at every iteration. To examine the effectiveness of our methods, we develop the benchmark suite of Safe Multi-Agent MuJoCo that involves a variety of MARL baselines. Experimental results justify that MACPO/MAPPO-Lagrangian can consistently satisfy safety constraints, meanwhile achieving comparable performance to strong baselines.","url_abs":"https://arxiv.org/abs/2110.02793v2","url_pdf":"https://arxiv.org/pdf/2110.02793v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"multi-agent-constrained-policy-optimisation","repo_url":"https://github.com/chauncygu/multi-agent-constrained-policy-optimisation","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"multi-agent-constrained-policy-optimisation","repo_url":"https://github.com/chauncygu/safe-multi-agent-isaac-gym","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"Apache-2.0"}},{"paper_slug":"multi-agent-constrained-policy-optimisation","repo_url":"https://github.com/chauncygu/safe-multi-agent-mujoco","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"multi-agent-constrained-policy-optimisation","repo_url":"https://github.com/chauncygu/safe-multi-agent-robosuite","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"mujoco","task_name":"MuJoCo"},{"task_slug":"multi-agent-reinforcement-learning","task_name":"Multi-agent Reinforcement Learning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2110.02793","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2110.02793"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/chauncygu/safe-multi-agent-mujoco","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/chauncygu/safe-multi-agent-robosuite","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/chauncygu/multi-agent-constrained-policy-optimisation","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/chauncygu/safe-multi-agent-isaac-gym","reach":{"status":"ok","spdx":"Apache-2.0"}}],"summary":{"ran_draft_wrong":1,"unverified":6},"by_repo_kind":{"official":{"samples":1,"ran":0,"repositories":1},"listed":{"samples":6,"ran":1,"repositories":2}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":1,"samples":[{"code_sha256_prefix":"9d0e7c23576a092f","entry":"convert","repo":"chauncygu/safe-multi-agent-robosuite","repo_kind":"listed","path":"robosuite/multi_agent/manyrobot_env.py","file_url":"https://github.com/chauncygu/safe-multi-agent-robosuite/blob/HEAD/robosuite/multi_agent/manyrobot_env.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"9d0e7c23576a092f"}},{"code_sha256_prefix":"7b68f30bd920cce1","entry":"build_obs","repo":"chauncygu/safe-multi-agent-mujoco","repo_kind":"listed","path":"safety_multi_agent_mujoco/safety_ma_mujoco/safety_multiagent_mujoco/obsk.py","file_url":"https://github.com/chauncygu/safe-multi-agent-mujoco/blob/HEAD/safety_multi_agent_mujoco/safety_ma_mujoco/safety_multiagent_mujoco/obsk.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"7b68f30bd920cce1"}},{"code_sha256_prefix":"c81e179f5e2f405f","entry":"convert_observation_to_space","repo":"chauncygu/safe-multi-agent-mujoco","repo_kind":"listed","path":"safety_multi_agent_mujoco/safety_ma_mujoco/safety_multiagent_mujoco/mujoco_env.py","file_url":"https://github.com/chauncygu/safe-multi-agent-mujoco/blob/HEAD/safety_multi_agent_mujoco/safety_ma_mujoco/safety_multiagent_mujoco/mujoco_env.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"c81e179f5e2f405f"}},{"code_sha256_prefix":"b740cb6f99054211","entry":"get_joints_at_kdist","repo":"chauncygu/safe-multi-agent-mujoco","repo_kind":"listed","path":"safety_multi_agent_mujoco/safety_ma_mujoco/safety_multiagent_mujoco/obsk.py","file_url":"https://github.com/chauncygu/safe-multi-agent-mujoco/blob/HEAD/safety_multi_agent_mujoco/safety_ma_mujoco/safety_multiagent_mujoco/obsk.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"b740cb6f99054211"}},{"code_sha256_prefix":"0e771fe439b0e7f9","entry":"get_parts_and_edges","repo":"chauncygu/safe-multi-agent-mujoco","repo_kind":"listed","path":"safety_multi_agent_mujoco/safety_ma_mujoco/safety_multiagent_mujoco/obsk.py","file_url":"https://github.com/chauncygu/safe-multi-agent-mujoco/blob/HEAD/safety_multi_agent_mujoco/safety_ma_mujoco/safety_multiagent_mujoco/obsk.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"0e771fe439b0e7f9"}},{"code_sha256_prefix":"46bd205700c42f20","entry":"mass_center","repo":"chauncygu/safe-multi-agent-mujoco","repo_kind":"listed","path":"safety_multi_agent_mujoco/safety_ma_mujoco/safety_multiagent_mujoco/humanoid.py","file_url":"https://github.com/chauncygu/safe-multi-agent-mujoco/blob/HEAD/safety_multi_agent_mujoco/safety_ma_mujoco/safety_multiagent_mujoco/humanoid.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"46bd205700c42f20"}},{"code_sha256_prefix":"5c01d0fdf71a654e","entry":"parse_args","repo":"chauncygu/multi-agent-constrained-policy-optimisation","repo_kind":"official","path":"MACPO/macpo/scripts/train/train_mujoco.py","file_url":"https://github.com/chauncygu/multi-agent-constrained-policy-optimisation/blob/HEAD/MACPO/macpo/scripts/train/train_mujoco.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"5c01d0fdf71a654e"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}