{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/parametrized-deep-q-networks-learning-playing","title":"PARAMETRIZED DEEP Q-NETWORKS LEARNING: PLAYING ONLINE BATTLE ARENA WITH DISCRETE-CONTINUOUS HYBRID ACTION SPACE","arxiv_id":null,"date":"2018-01-01","proceeding":"ICLR 2018 1","authors":["Jiechao Xiong","Qing Wang","Zhuoran Yang","Peng Sun","Yang Zheng","Lei Han","Haobo Fu","Xiangru Lian","Carson Eisenach","Haichuan Yang","Emmanuel Ekwedike","Bei Peng","Haoyue Gao","Tong Zhang","Ji Liu","Han Liu"],"abstract":"Most existing deep reinforcement learning (DRL) frameworks consider action spaces that are either\ndiscrete or continuous space. Motivated by the project of design Game AI for King of Glory\n(KOG), one the world’s most popular mobile game, we consider the scenario with the discrete-continuous\nhybrid action space. To directly apply existing DLR frameworks, existing approaches\neither approximate the hybrid space by a discrete set or relaxing it into a continuous set, which is\nusually less efficient and robust. In this paper, we propose a parametrized deep Q-network (P-DQN)\nfor the hybrid action space without approximation or relaxation. Our algorithm combines DQN and\nDDPG and can be viewed as an extension of the DQN to hybrid actions. The empirical study on the\ngame KOG validates the efficiency and effectiveness of our method.","url_abs":"https://openreview.net/forum?id=Sy_MK3lAZ","url_pdf":"https://openreview.net/pdf?id=Sy_MK3lAZ","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"parametrized-deep-q-networks-learning-playing","repo_url":"https://github.com/opendilab/DI-engine","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"deep-reinforcement-learning","task_name":"Deep Reinforcement Learning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"}],"methods":[{"method_slug":"convolution","method_name":"Convolution"},{"method_slug":"dqn","method_name":"DQN"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"q-learning","method_name":"Q-Learning"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}