{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/evaluating-multi-agent-coordination-abilities","title":"LLM-Coordination: Evaluating and Analyzing Multi-agent Coordination Abilities in Large Language Models","arxiv_id":"2310.03903","date":"2023-10-05","proceeding":null,"authors":["Saaket Agashe","Yue Fan","Anthony Reyna","Xin Eric Wang"],"abstract":"Large Language Models (LLMs) have demonstrated emergent common-sense reasoning and Theory of Mind (ToM) capabilities, making them promising candidates for developing coordination agents. This study introduces the LLM-Coordination Benchmark, a novel benchmark for analyzing LLMs in the context of Pure Coordination Settings, where agents must cooperate to maximize gains. Our benchmark evaluates LLMs through two distinct tasks. The first is Agentic Coordination, where LLMs act as proactive participants in four pure coordination games. The second is Coordination Question Answering (CoordQA), which tests LLMs on 198 multiple-choice questions across these games to evaluate three key abilities: Environment Comprehension, ToM Reasoning, and Joint Planning. Results from Agentic Coordination experiments reveal that LLM-Agents excel in multi-agent coordination settings where decision-making primarily relies on environmental variables but face challenges in scenarios requiring active consideration of partners' beliefs and intentions. The CoordQA experiments further highlight significant room for improvement in LLMs' Theory of Mind reasoning and joint planning capabilities. Zero-Shot Coordination (ZSC) experiments in the Agentic Coordination setting demonstrate that LLM agents, unlike RL methods, exhibit robustness to unseen partners. These findings indicate the potential of LLMs as Agents in pure coordination setups and underscore areas for improvement. Code Available at https://github.com/eric-ai-lab/llm_coordination.","url_abs":"https://arxiv.org/abs/2310.03903v3","url_pdf":"https://arxiv.org/pdf/2310.03903v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"evaluating-multi-agent-coordination-abilities","repo_url":"https://github.com/eric-ai-lab/llm_coordination","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"common-sense-reasoning","task_name":"Common Sense Reasoning"},{"task_slug":"multiple-choice","task_name":"Multiple-choice"},{"task_slug":"question-answering","task_name":"Question Answering"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2310.03903","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2310.03903"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/eric-ai-lab/llm_coordination","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"ran":4,"unverified":2},"by_repo_kind":{"official":{"samples":6,"ran":4,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"d5d2bac74a982663","entry":"ai_score","repo":"eric-ai-lab/llm_coordination","repo_kind":"official","path":"src/llm_coordination_agents/hanabi_action_manager.py","file_url":"https://github.com/eric-ai-lab/llm_coordination/blob/HEAD/src/llm_coordination_agents/hanabi_action_manager.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"d5d2bac74a982663"}},{"code_sha256_prefix":"580ac70c7760b12d","entry":"dict_from_response","repo":"eric-ai-lab/llm_coordination","repo_kind":"official","path":"hanabi_bot/commands_websocket.py","file_url":"https://github.com/eric-ai-lab/llm_coordination/blob/HEAD/hanabi_bot/commands_websocket.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"580ac70c7760b12d"}},{"code_sha256_prefix":"76f68dc70352e2e8","entry":"extract_location","repo":"eric-ai-lab/llm_coordination","repo_kind":"official","path":"src/llm_coordination_agents/overcooked_action_manager.py","file_url":"https://github.com/eric-ai-lab/llm_coordination/blob/HEAD/src/llm_coordination_agents/overcooked_action_manager.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"76f68dc70352e2e8"}},{"code_sha256_prefix":"03c14009418b699b","entry":"gameJoin","repo":"eric-ai-lab/llm_coordination","repo_kind":"official","path":"hanabi_bot/commands_websocket.py","file_url":"https://github.com/eric-ai-lab/llm_coordination/blob/HEAD/hanabi_bot/commands_websocket.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"03c14009418b699b"}},{"code_sha256_prefix":"3a41be75f86be0ff","entry":"gameCreate","repo":"eric-ai-lab/llm_coordination","repo_kind":"official","path":"hanabi_bot/commands_websocket.py","file_url":"https://github.com/eric-ai-lab/llm_coordination/blob/HEAD/hanabi_bot/commands_websocket.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"3a41be75f86be0ff"}},{"code_sha256_prefix":"a801f954b19cff10","entry":"get_user_action","repo":"eric-ai-lab/llm_coordination","repo_kind":"official","path":"src/llm_coordination_agents/collab_capture_action_manager.py","file_url":"https://github.com/eric-ai-lab/llm_coordination/blob/HEAD/src/llm_coordination_agents/collab_capture_action_manager.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"a801f954b19cff10"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}