{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/enhancing-cooperative-multi-agent","title":"Enhancing Cooperative Multi-Agent Reinforcement Learning with State Modelling and Adversarial Exploration","arxiv_id":"2505.05262","date":"2025-05-08","proceeding":null,"authors":["Andreas Kontogiannis","Konstantinos Papathanasiou","Yi Shen","Giorgos Stamou","Michael M. Zavlanos","George Vouros"],"abstract":"Learning to cooperate in distributed partially observable environments with no communication abilities poses significant challenges for multi-agent deep reinforcement learning (MARL). This paper addresses key concerns in this domain, focusing on inferring state representations from individual agent observations and leveraging these representations to enhance agents' exploration and collaborative task execution policies. To this end, we propose a novel state modelling framework for cooperative MARL, where agents infer meaningful belief representations of the non-observable state, with respect to optimizing their own policies, while filtering redundant and less informative joint state information. Building upon this framework, we propose the MARL SMPE algorithm. In SMPE, agents enhance their own policy's discriminative abilities under partial observability, explicitly by incorporating their beliefs into the policy network, and implicitly by adopting an adversarial type of exploration policies which encourages agents to discover novel, high-value states while improving the discriminative abilities of others. Experimentally, we show that SMPE outperforms state-of-the-art MARL algorithms in complex fully cooperative tasks from the MPE, LBF, and RWARE benchmarks.","url_abs":"https://arxiv.org/abs/2505.05262v2","url_pdf":"https://arxiv.org/pdf/2505.05262v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"enhancing-cooperative-multi-agent","repo_url":"https://github.com/ddaedalus/smpe","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"deep-reinforcement-learning","task_name":"Deep Reinforcement Learning"},{"task_slug":"multi-agent-reinforcement-learning","task_name":"Multi-agent Reinforcement Learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2505.05262","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2505.05262"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"deterministic:regex_extraction","url":"https://github.com/uoe-agents/lb-foraging","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"deterministic:regex_extraction","url":"https://github.com/semitable/multiagent-particle-envs","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"deterministic:regex_extraction","url":"https://github.com/uoe-agents/robotic-warehouse","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/ddaedalus/smpe","reach":null}],"summary":{"ran":3,"ran_honours":1,"unverified":6},"by_repo_kind":{"official":{"samples":4,"ran":4,"repositories":1},"found_in_text":{"samples":6,"ran":0,"repositories":3}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":4,"samples":[{"code_sha256_prefix":"249f642354cef321","entry":"Decoder","repo":"ddaedalus/smpe","repo_kind":"official","path":"modules/dynamics/variational_inference.py","file_url":"https://github.com/ddaedalus/smpe/blob/HEAD/modules/dynamics/variational_inference.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"249f642354cef321"}},{"code_sha256_prefix":"6899142d57307b36","entry":"VAE","repo":"ddaedalus/smpe","repo_kind":"official","path":"modules/dynamics/variational_inference.py","file_url":"https://github.com/ddaedalus/smpe/blob/HEAD/modules/dynamics/variational_inference.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"6899142d57307b36"}},{"code_sha256_prefix":"cbb8702b81639880","entry":"VariationalEncoder","repo":"ddaedalus/smpe","repo_kind":"official","path":"modules/dynamics/variational_inference.py","file_url":"https://github.com/ddaedalus/smpe/blob/HEAD/modules/dynamics/variational_inference.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"cbb8702b81639880"}},{"code_sha256_prefix":"9236867bb372c687","entry":"kl_distance","repo":"ddaedalus/smpe","repo_kind":"official","path":"modules/dynamics/variational_inference.py","file_url":"https://github.com/ddaedalus/smpe/blob/HEAD/modules/dynamics/variational_inference.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"9236867bb372c687"}},{"code_sha256_prefix":"10da81242bba7f38","entry":"get_display","repo":"semitable/multiagent-particle-envs","repo_kind":"found_in_text","path":"mpe/rendering.py","file_url":"https://github.com/semitable/multiagent-particle-envs/blob/HEAD/mpe/rendering.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":false,"mcp_get_code":{"code_sha256":"10da81242bba7f38"}},{"code_sha256_prefix":"ac12233b3f8cddcf","entry":"get_display","repo":"uoe-agents/lb-foraging","repo_kind":"found_in_text","path":"lbforaging/foraging/rendering.py","file_url":"https://github.com/uoe-agents/lb-foraging/blob/HEAD/lbforaging/foraging/rendering.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"ac12233b3f8cddcf"}},{"code_sha256_prefix":"d681f8b5572da32a","entry":"get_display","repo":"uoe-agents/robotic-warehouse","repo_kind":"found_in_text","path":"rware/rendering.py","file_url":"https://github.com/uoe-agents/robotic-warehouse/blob/HEAD/rware/rendering.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"d681f8b5572da32a"}},{"code_sha256_prefix":"f8dd361d2679d007","entry":"make_circle","repo":"semitable/multiagent-particle-envs","repo_kind":"found_in_text","path":"mpe/rendering.py","file_url":"https://github.com/semitable/multiagent-particle-envs/blob/HEAD/mpe/rendering.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":false,"mcp_get_code":{"code_sha256":"f8dd361d2679d007"}},{"code_sha256_prefix":"e0db348c4be68fc5","entry":"make_env","repo":"semitable/multiagent-particle-envs","repo_kind":"found_in_text","path":"make_env.py","file_url":"https://github.com/semitable/multiagent-particle-envs/blob/HEAD/make_env.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":false,"mcp_get_code":{"code_sha256":"e0db348c4be68fc5"}},{"code_sha256_prefix":"ea4edded084fb348","entry":"make_polygon","repo":"semitable/multiagent-particle-envs","repo_kind":"found_in_text","path":"mpe/rendering.py","file_url":"https://github.com/semitable/multiagent-particle-envs/blob/HEAD/mpe/rendering.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":false,"mcp_get_code":{"code_sha256":"ea4edded084fb348"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}