{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/modeling-others-using-oneself-in-multi-agent","title":"Modeling Others using Oneself in Multi-Agent Reinforcement Learning","arxiv_id":"1802.09640","date":"2018-02-26","proceeding":"ICML 2018 7","authors":["Roberta Raileanu","Emily Denton","Arthur Szlam","Rob Fergus"],"abstract":"We consider the multi-agent reinforcement learning setting with imperfect\ninformation in which each agent is trying to maximize its own utility. The\nreward function depends on the hidden state (or goal) of both agents, so the\nagents must infer the other players' hidden goals from their observed behavior\nin order to solve the tasks. We propose a new approach for learning in these\ndomains: Self Other-Modeling (SOM), in which an agent uses its own policy to\npredict the other agent's actions and update its belief of their hidden state\nin an online manner. We evaluate this approach on three different tasks and\nshow that the agents are able to learn better policies using their estimate of\nthe other players' hidden states, in both cooperative and adversarial settings.","url_abs":"http://arxiv.org/abs/1802.09640v3","url_pdf":"http://arxiv.org/pdf/1802.09640v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"modeling-others-using-oneself-in-multi-agent","repo_url":"https://github.com/cts198859/deeprl_dist","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}}],"tasks":[{"task_slug":"multi-agent-reinforcement-learning","task_name":"Multi-agent Reinforcement Learning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1802.09640","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}