{"url":"/method/dqn","slug":"dqn","name":"DQN","full_name":"Deep Q-Network","full_name_withheld":false,"description_markdown":"A **DQN**, or Deep Q-Network, approximates a state-value function in a [Q-Learning](https://paperswithcode.com/method/q-learning) framework with a neural network. In the Atari Games case, they take in several frames of the game as an input and output state values for each action as an output. \r\n\r\nIt is usually used in conjunction with [Experience Replay](https://paperswithcode.com/method/experience-replay), for storing the episode steps in memory for off-policy learning, where samples are drawn from the replay memory at random. Additionally, the Q-Network is usually optimized towards a frozen target network that is periodically updated with the latest weights every $k$ steps (where $k$ is a hyperparameter). The latter makes training more stable by preventing short-term oscillations from a moving target. The former tackles autocorrelation that would occur from on-line learning, and having a replay memory makes the problem more like a supervised learning problem.\r\n\r\nImage Source: [here](https://www.researchgate.net/publication/319643003_Autonomous_Quadrotor_Landing_using_Deep_Reinforcement_Learning)","description_state":"present","introduced_year":null,"introduced_by":{"title":"Playing Atari with Deep Reinforcement Learning","paper":"/paper/playing-atari-with-deep-reinforcement","first_author":"Volodymyr Mnih","n_authors":7,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/playing-atari-with-deep-reinforcement"},"source":{"url":"http://arxiv.org/abs/1312.5602v1","title":"Playing Atari with Deep Reinforcement Learning","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Reinforcement Learning","area_id":"reinforcement-learning","collection":"Q-Learning Networks","url":"/methods/category/q-learning-networks","pwc_aliases":["q-learning"]}],"n_papers_tagged":519,"archive_num_papers":519,"papers_newest_first":[{"paper":null,"title":"Turning Sand to Gold: Recycling Data to Bridge On-Policy and Off-Policy Learning via Causal Bound","date":"2025-07-15","arxiv_id":"2507.11269","n_code_links":0,"syntology":null},{"paper":null,"title":"Detecting and Mitigating Reward Hacking in Reinforcement Learning Systems: A Comprehensive Empirical Study","date":"2025-07-08","arxiv_id":"2507.05619","n_code_links":0,"syntology":null},{"paper":null,"title":"2048: Reinforcement Learning in a Delayed Reward Environment","date":"2025-07-07","arxiv_id":"2507.05465","n_code_links":0,"syntology":null},{"paper":null,"title":"VRAIL: Vectorized Reward-based Attribution for Interpretable Learning","date":"2025-06-19","arxiv_id":"2506.16014","n_code_links":0,"syntology":null},{"paper":null,"title":"GCN-Driven Reinforcement Learning for Probabilistic Real-Time Guarantees in Industrial URLLC","date":"2025-06-17","arxiv_id":"2506.15011","n_code_links":0,"syntology":null},{"paper":null,"title":"Reliable Critics: Monotonic Improvement and Convergence Guarantees for Reinforcement Learning","date":"2025-06-08","arxiv_id":"2506.07134","n_code_links":0,"syntology":null},{"paper":null,"title":"Getting More from Less: Transfer Learning Improves Sleep Stage Decoding Accuracy in Peripheral Wearable Devices","date":"2025-05-31","arxiv_id":"2506.00730","n_code_links":0,"syntology":null},{"paper":null,"title":"Combining Deep Architectures for Information Gain estimation and Reinforcement Learning for multiagent field exploration","date":"2025-05-29","arxiv_id":"2505.23865","n_code_links":0,"syntology":null},{"paper":"/paper/the-cell-must-go-on-agar-io-for-continual","title":"The Cell Must Go On: Agar.io for Continual Reinforcement Learning","date":"2025-05-23","arxiv_id":"2505.18347","n_code_links":1,"syntology":null},{"paper":null,"title":"LLM-Explorer: A Plug-in Reinforcement Learning Policy Exploration Enhancement Driven by Large Language Models","date":"2025-05-21","arxiv_id":"2505.15293","n_code_links":0,"syntology":null},{"paper":null,"title":"Automatic Reward Shaping from Confounded Offline Data","date":"2025-05-16","arxiv_id":"2505.11478","n_code_links":0,"syntology":null},{"paper":null,"title":"Reinforcement Learning for Game-Theoretic Resource Allocation on Graphs","date":"2025-05-08","arxiv_id":"2505.06319","n_code_links":0,"syntology":null},{"paper":null,"title":"Interpretable Learning Dynamics in Unsupervised Reinforcement Learning","date":"2025-05-06","arxiv_id":"2505.06279","n_code_links":0,"syntology":null},{"paper":null,"title":"Universal Approximation Theorem of Deep Q-Networks","date":"2025-05-04","arxiv_id":"2505.02288","n_code_links":0,"syntology":null},{"paper":null,"title":"Approximation to Deep Q-Network by Stochastic Delay Differential Equations","date":"2025-05-01","arxiv_id":"2505.00382","n_code_links":0,"syntology":null},{"paper":null,"title":"AlphaGrad: Non-Linear Gradient Normalization Optimizer","date":"2025-04-22","arxiv_id":"2504.16020","n_code_links":0,"syntology":null},{"paper":null,"title":"State-Aware IoT Scheduling Using Deep Q-Networks and Edge-Based Coordination","date":"2025-04-22","arxiv_id":"2504.15577","n_code_links":0,"syntology":null},{"paper":null,"title":"Graph Based Deep Reinforcement Learning Aided by Transformers for Multi-Agent Cooperation","date":"2025-04-11","arxiv_id":"2504.08195","n_code_links":0,"syntology":null},{"paper":null,"title":"Dynamic Operating System Scheduling Using Double DQN: A Reinforcement Learning Approach to Task Optimization","date":"2025-03-31","arxiv_id":"2503.23659","n_code_links":0,"syntology":null},{"paper":null,"title":"Deep Q-Learning with Gradient Target Tracking","date":"2025-03-20","arxiv_id":"2503.16700","n_code_links":0,"syntology":null},{"paper":"/paper/a-generalist-hanabi-agent","title":"A Generalist Hanabi Agent","date":"2025-03-17","arxiv_id":"2503.14555","n_code_links":1,"syntology":null},{"paper":null,"title":"Exploring Competitive and Collusive Behaviors in Algorithmic Pricing with Deep Reinforcement Learning","date":"2025-03-14","arxiv_id":"2503.11270","n_code_links":0,"syntology":null},{"paper":null,"title":"Intelligent Joint Security and Delay Determinacy Performance Guarantee Strategy in RIS-Assisted IIoT Communication Systems","date":"2025-03-11","arxiv_id":"2503.08086","n_code_links":0,"syntology":null},{"paper":null,"title":"Human-AI Collaboration in Cloud Security: Cognitive Hierarchy-Driven Deep Reinforcement Learning","date":"2025-02-22","arxiv_id":"2502.16054","n_code_links":0,"syntology":null},{"paper":"/paper/ranking-joint-policies-in-dynamic-games-using","title":"Ranking Joint Policies in Dynamic Games using Evolutionary Dynamics","date":"2025-02-20","arxiv_id":"2502.14724","n_code_links":1,"syntology":null},{"paper":null,"title":"Seasonal Station-Keeping of Short Duration High Altitude Balloons using Deep Reinforcement Learning","date":"2025-02-07","arxiv_id":"2502.05014","n_code_links":0,"syntology":null},{"paper":null,"title":"Reinforcement Learning for Quantum Circuit Design: Using Matrix Representations","date":"2025-01-27","arxiv_id":"2501.16509","n_code_links":0,"syntology":null},{"paper":null,"title":"Optimizing Return Distributions with Distributional Dynamic Programming","date":"2025-01-22","arxiv_id":"2501.13028","n_code_links":0,"syntology":null},{"paper":null,"title":"Perception-Guided EEG Analysis: A Deep Learning Approach Inspired by Level of Detail (LOD) Theory","date":"2025-01-11","arxiv_id":"2501.10428","n_code_links":0,"syntology":null},{"paper":null,"title":"Session-Level Dynamic Ad Load Optimization using Offline Robust Reinforcement Learning","date":"2025-01-09","arxiv_id":"2501.05591","n_code_links":0,"syntology":null}],"papers_shown":30,"tasks":[{"task":"/task/reinforcement-learning-1","name":"Reinforcement Learning (RL)","papers":299},{"task":"/task/reinforcement-learning","name":"Reinforcement Learning","papers":265},{"task":"/task/reinforcement-learning-2","name":"reinforcement-learning","papers":260},{"task":"/task/deep-reinforcement-learning","name":"Deep Reinforcement Learning","papers":217},{"task":"/task/q-learning","name":"Q-Learning","papers":116},{"task":"/task/atari-games","name":"Atari Games","papers":68},{"task":"/task/decision-making","name":"Decision Making","papers":38},{"task":"/task/management","name":"Management","papers":21},{"task":"/task/efficient-exploration","name":"Efficient Exploration","papers":17},{"task":"/task/multi-agent-reinforcement-learning","name":"Multi-agent Reinforcement Learning","papers":17},{"task":"/task/openai-gym","name":"OpenAI Gym","papers":14},{"task":"/task/scheduling","name":"Scheduling","papers":14},{"task":"/task/autonomous-driving","name":"Autonomous Driving","papers":13},{"task":"/task/continuous-control","name":"Continuous Control","papers":10},{"task":"/task/diversity","name":"Diversity","papers":9},{"task":"/task/transfer-learning","name":"Transfer Learning","papers":9},{"task":"/task/continuous-control","name":"continuous-control","papers":9},{"task":"/task/benchmarking","name":"Benchmarking","papers":8},{"task":"/task/eeg-1","name":"EEG","papers":8},{"task":"/task/imitation-learning","name":"Imitation Learning","papers":8}],"tasks_shown":20,"n_tasks":244,"usage_by_year":[{"year":"2013","papers":1},{"year":"2014","papers":1},{"year":"2015","papers":8},{"year":"2016","papers":9},{"year":"2017","papers":24},{"year":"2018","papers":40},{"year":"2019","papers":52},{"year":"2020","papers":72},{"year":"2021","papers":92},{"year":"2022","papers":53},{"year":"2023","papers":70},{"year":"2024","papers":65},{"year":"2025","papers":32}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/dqn"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}