{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/competitive-multi-agent-inverse-reinforcement","title":"Competitive Multi-agent Inverse Reinforcement Learning with Sub-optimal Demonstrations","arxiv_id":"1801.02124","date":"2018-01-07","proceeding":"ICML 2018 7","authors":["Xingyu Wang","Diego Klabjan"],"abstract":"This paper considers the problem of inverse reinforcement learning in\nzero-sum stochastic games when expert demonstrations are known to be not\noptimal. Compared to previous works that decouple agents in the game by\nassuming optimality in expert strategies, we introduce a new objective function\nthat directly pits experts against Nash Equilibrium strategies, and we design\nan algorithm to solve for the reward function in the context of inverse\nreinforcement learning with deep neural networks as model approximations. In\nour setting the model and algorithm do not decouple by agent. In order to find\nNash Equilibrium in large-scale games, we also propose an adversarial training\nalgorithm for zero-sum stochastic games, and show the theoretical appeal of\nnon-existence of local optima in its objective function. In our numerical\nexperiments, we demonstrate that our Nash Equilibrium and inverse reinforcement\nlearning algorithms address games that are not amenable to previous approaches\nusing tabular representations. Moreover, with sub-optimal expert demonstrations\nour algorithms recover both reward functions and strategies with good quality.","url_abs":"http://arxiv.org/abs/1801.02124v2","url_pdf":"http://arxiv.org/pdf/1801.02124v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"competitive-multi-agent-inverse-reinforcement","repo_url":"https://github.com/mdabbah/pacman-rl-project","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1801.02124","atlas_url":"https://app.syntology.ai/?focus=1801.02124","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}