{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/robust-multi-agent-reinforcement-learning-3","title":"Robust Multi-Agent Reinforcement Learning with State Uncertainty","arxiv_id":"2307.16212","date":"2023-07-30","proceeding":null,"authors":["Sihong He","Songyang Han","Sanbao Su","Shuo Han","Shaofeng Zou","Fei Miao"],"abstract":"In real-world multi-agent reinforcement learning (MARL) applications, agents may not have perfect state information (e.g., due to inaccurate measurement or malicious attacks), which challenges the robustness of agents' policies. Though robustness is getting important in MARL deployment, little prior work has studied state uncertainties in MARL, neither in problem formulation nor algorithm design. Motivated by this robustness issue and the lack of corresponding studies, we study the problem of MARL with state uncertainty in this work. We provide the first attempt to the theoretical and empirical analysis of this challenging problem. We first model the problem as a Markov Game with state perturbation adversaries (MG-SPA) by introducing a set of state perturbation adversaries into a Markov Game. We then introduce robust equilibrium (RE) as the solution concept of an MG-SPA. We conduct a fundamental analysis regarding MG-SPA such as giving conditions under which such a robust equilibrium exists. Then we propose a robust multi-agent Q-learning (RMAQ) algorithm to find such an equilibrium, with convergence guarantees. To handle high-dimensional state-action space, we design a robust multi-agent actor-critic (RMAAC) algorithm based on an analytical expression of the policy gradient derived in the paper. Our experiments show that the proposed RMAQ algorithm converges to the optimal value function; our RMAAC algorithm outperforms several MARL and robust MARL methods in multiple multi-agent environments when state uncertainty is present. The source code is public on \\url{https://github.com/sihongho/robust_marl_with_state_uncertainty}.","url_abs":"https://arxiv.org/abs/2307.16212v1","url_pdf":"https://arxiv.org/pdf/2307.16212v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"robust-multi-agent-reinforcement-learning-3","repo_url":"https://github.com/sihongho/robust_marl_with_state_uncertainty","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"tf","reach":{"status":"ok"}}],"tasks":[{"task_slug":"multi-agent-reinforcement-learning","task_name":"Multi-agent Reinforcement Learning"},{"task_slug":"q-learning","task_name":"Q-Learning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"single-particle-analysis","task_name":"Single Particle Analysis"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[{"method_slug":"q-learning","method_name":"Q-Learning"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2307.16212","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2307.16212"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/sihongho/robust_marl_with_state_uncertainty","reach":{"status":"ok"}}],"summary":{"ran_violates":1,"ran":1,"unverified":3},"by_repo_kind":{"official":{"samples":5,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":5,"samples":[{"code_sha256_prefix":"fb985d688435bf65","entry":"discount_with_dones","repo":"sihongho/robust_marl_with_state_uncertainty","repo_kind":"official","path":"RMA-AC/maddpg/trainer/maddpg.py","file_url":"https://github.com/sihongho/robust_marl_with_state_uncertainty/blob/HEAD/RMA-AC/maddpg/trainer/maddpg.py","link_basis":"plan_row","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":"well_formed","behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"fb985d688435bf65"}},{"code_sha256_prefix":"aad5113d0b095522","entry":"shape_el","repo":"sihongho/robust_marl_with_state_uncertainty","repo_kind":"official","path":"RMA-AC/maddpg/common/distributions.py","file_url":"https://github.com/sihongho/robust_marl_with_state_uncertainty/blob/HEAD/RMA-AC/maddpg/common/distributions.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"aad5113d0b095522"}},{"code_sha256_prefix":"8ebadf88476f8efe","entry":"mean","repo":"sihongho/robust_marl_with_state_uncertainty","repo_kind":"official","path":"RMA-AC/maddpg/common/tf_util.py","file_url":"https://github.com/sihongho/robust_marl_with_state_uncertainty/blob/HEAD/RMA-AC/maddpg/common/tf_util.py","link_basis":"plan_row","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"8ebadf88476f8efe"}},{"code_sha256_prefix":"fbab233112bf4d1e","entry":"sum","repo":"sihongho/robust_marl_with_state_uncertainty","repo_kind":"official","path":"RMA-AC/maddpg/common/tf_util.py","file_url":"https://github.com/sihongho/robust_marl_with_state_uncertainty/blob/HEAD/RMA-AC/maddpg/common/tf_util.py","link_basis":"plan_row","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"fbab233112bf4d1e"}},{"code_sha256_prefix":"598184a2689d36d0","entry":"var","repo":"sihongho/robust_marl_with_state_uncertainty","repo_kind":"official","path":"RMA-AC/maddpg/common/tf_util.py","file_url":"https://github.com/sihongho/robust_marl_with_state_uncertainty/blob/HEAD/RMA-AC/maddpg/common/tf_util.py","link_basis":"plan_row","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"598184a2689d36d0"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}