{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/opponent-modeling-in-deep-reinforcement","title":"Opponent Modeling in Deep Reinforcement Learning","arxiv_id":"1609.05559","date":"2016-09-18","proceeding":null,"authors":["He He","Jordan Boyd-Graber","Kevin Kwok","Hal Daumé III"],"abstract":"Opponent modeling is necessary in multi-agent settings where secondary agents\nwith competing goals also adapt their strategies, yet it remains challenging\nbecause strategies interact with each other and change. Most previous work\nfocuses on developing probabilistic models or parameterized strategies for\nspecific applications. Inspired by the recent success of deep reinforcement\nlearning, we present neural-based models that jointly learn a policy and the\nbehavior of opponents. Instead of explicitly predicting the opponent's action,\nwe encode observation of the opponents into a deep Q-Network (DQN); however, we\nretain explicit modeling (if desired) using multitasking. By using a\nMixture-of-Experts architecture, our model automatically discovers different\nstrategy patterns of opponents without extra supervision. We evaluate our\nmodels on a simulated soccer game and a popular trivia game, showing superior\nperformance over DQN and its variants.","url_abs":"http://arxiv.org/abs/1609.05559v1","url_pdf":"http://arxiv.org/pdf/1609.05559v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"opponent-modeling-in-deep-reinforcement","repo_url":"https://github.com/hhexiy/opponent","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"torch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"deep-reinforcement-learning","task_name":"Deep Reinforcement Learning"},{"task_slug":"mixture-of-experts","task_name":"Mixture-of-Experts"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[{"method_slug":"convolution","method_name":"Convolution"},{"method_slug":"dqn","method_name":"DQN"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"q-learning","method_name":"Q-Learning"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1609.05559","atlas_url":"https://app.syntology.ai/?focus=1609.05559","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}