{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/combinational-q-learning-for-dou-di-zhu","title":"Combinational Q-Learning for Dou Di Zhu","arxiv_id":"1901.08925","date":"2019-01-24","proceeding":null,"authors":["Yang You","Liangwei Li","Baisong Guo","Weiming Wang","Cewu Lu"],"abstract":"Deep reinforcement learning (DRL) has gained a lot of attention in recent\nyears, and has been proven to be able to play Atari games and Go at or above\nhuman levels. However, those games are assumed to have a small fixed number of\nactions and could be trained with a simple CNN network. In this paper, we study\na special class of Asian popular card games called Dou Di Zhu, in which two\nadversarial groups of agents must consider numerous card combinations at each\ntime step, leading to huge number of actions. We propose a novel method to\nhandle combinatorial actions, which we call combinational Q-learning (CQL). We\nemploy a two-stage network to reduce action space and also leverage\norder-invariant max-pooling operations to extract relationships between\nprimitive actions. Results show that our method prevails over state-of-the art\nmethods like naive Q-learning and A3C. We develop an easy-to-use card game\nenvironments and train all agents adversarially from sractch, with only\nknowledge of game rules and verify that our agents are comparative to humans.\nOur code to reproduce all reported results will be available online.","url_abs":"http://arxiv.org/abs/1901.08925v2","url_pdf":"http://arxiv.org/pdf/1901.08925v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"combinational-q-learning-for-dou-di-zhu","repo_url":"https://github.com/qq456cvb/doudizhu-C","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}}],"tasks":[{"task_slug":"atari-games","task_name":"Atari Games"},{"task_slug":"card-games","task_name":"Card Games"},{"task_slug":"deep-reinforcement-learning","task_name":"Deep Reinforcement Learning"},{"task_slug":"q-learning","task_name":"Q-Learning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"}],"methods":[{"method_slug":"a3c","method_name":"A3C"},{"method_slug":"convolution","method_name":"Convolution"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"entropy-regularization","method_name":"Entropy Regularization"},{"method_slug":"q-learning","method_name":"Q-Learning"},{"method_slug":"softmax","method_name":"Softmax"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1901.08925","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}