{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/gan-q-learning","title":"GAN Q-learning","arxiv_id":"1805.04874","date":"2018-05-13","proceeding":null,"authors":["Thang Doan","Bogdan Mazoure","Clare Lyle"],"abstract":"Distributional reinforcement learning (distributional RL) has seen empirical\nsuccess in complex Markov Decision Processes (MDPs) in the setting of nonlinear\nfunction approximation. However, there are many different ways in which one can\nleverage the distributional approach to reinforcement learning. In this paper,\nwe propose GAN Q-learning, a novel distributional RL method based on generative\nadversarial networks (GANs) and analyze its performance in simple tabular\nenvironments, as well as OpenAI Gym. We empirically show that our algorithm\nleverages the flexibility and blackbox approach of deep learning models while\nproviding a viable alternative to traditional methods.","url_abs":"http://arxiv.org/abs/1805.04874v3","url_pdf":"http://arxiv.org/pdf/1805.04874v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"gan-q-learning","repo_url":"https://github.com/daggertye/GAN-Q-Learning","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}}],"tasks":[{"task_slug":"distributional-reinforcement-learning","task_name":"Distributional Reinforcement Learning"},{"task_slug":"openai-gym","task_name":"OpenAI Gym"},{"task_slug":"q-learning","task_name":"Q-Learning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[{"method_slug":"convolution","method_name":"Convolution"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1805.04874","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}