{"url":"/method/ddpg","slug":"ddpg","name":"DDPG","full_name":"Deep Deterministic Policy Gradient","full_name_withheld":false,"description_markdown":"**DDPG**, or **Deep Deterministic Policy Gradient**, is an actor-critic, model-free algorithm based on the deterministic policy gradient that can operate over continuous action spaces. It combines the actor-critic approach with insights from [DQNs](https://paperswithcode.com/method/dqn): in particular, the insights that 1) the network is trained off-policy with samples from a replay buffer to minimize correlations between samples, and 2) the network is trained with a target Q network to give consistent targets during temporal difference backups. DDPG makes use of the same ideas along with [batch normalization](https://paperswithcode.com/method/batch-normalization).","description_state":"present","introduced_year":null,"introduced_by":{"title":"Continuous control with deep reinforcement learning","paper":"/paper/continuous-control-with-deep-reinforcement","first_author":"Timothy P. Lillicrap","n_authors":8,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/continuous-control-with-deep-reinforcement"},"source":{"url":"https://arxiv.org/abs/1509.02971v6","title":"Continuous control with deep reinforcement learning","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Reinforcement Learning","area_id":"reinforcement-learning","collection":"Policy Gradient Methods","url":"/methods/category/policy-gradient-methods","pwc_aliases":[]}],"n_papers_tagged":218,"archive_num_papers":218,"papers_newest_first":[{"paper":null,"title":"Multi-Objective Reinforcement Learning for Cognitive Radar Resource Management","date":"2025-06-25","arxiv_id":"2506.20853","n_code_links":0,"syntology":null},{"paper":null,"title":"Reliable Critics: Monotonic Improvement and Convergence Guarantees for Reinforcement Learning","date":"2025-06-08","arxiv_id":"2506.07134","n_code_links":0,"syntology":null},{"paper":null,"title":"A Novel Deep Reinforcement Learning Method for Computation Offloading in Multi-User Mobile Edge Computing with Decentralization","date":"2025-06-03","arxiv_id":"2506.02458","n_code_links":0,"syntology":null},{"paper":null,"title":"LLM-Explorer: A Plug-in Reinforcement Learning Policy Exploration Enhancement Driven by Large Language Models","date":"2025-05-21","arxiv_id":"2505.15293","n_code_links":0,"syntology":null},{"paper":null,"title":"Deep reinforcement learning-based longitudinal control strategy for automated vehicles at signalised intersections","date":"2025-05-13","arxiv_id":"2505.08896","n_code_links":0,"syntology":null},{"paper":null,"title":"Moderate Actor-Critic Methods: Controlling Overestimation Bias via Expectile Loss","date":"2025-04-14","arxiv_id":"2504.09929","n_code_links":0,"syntology":null},{"paper":null,"title":"Intelligent Joint Security and Delay Determinacy Performance Guarantee Strategy in RIS-Assisted IIoT Communication Systems","date":"2025-03-11","arxiv_id":"2503.08086","n_code_links":0,"syntology":null},{"paper":"/paper/evorl-a-gpu-accelerated-framework-for","title":"EvoRL: A GPU-accelerated Framework for Evolutionary Reinforcement Learning","date":"2025-01-25","arxiv_id":"2501.15129","n_code_links":2,"syntology":null},{"paper":null,"title":"Federated Deep Reinforcement Learning for Energy Efficient Multi-Functional RIS-Assisted Low-Earth Orbit Networks","date":"2025-01-19","arxiv_id":"2501.11079","n_code_links":0,"syntology":null},{"paper":null,"title":"AutoLoop: Fast Visual SLAM Fine-tuning through Agentic Curriculum Learning","date":"2025-01-15","arxiv_id":"2501.09160","n_code_links":0,"syntology":null},{"paper":null,"title":"Dynamic Portfolio Optimization via Augmented DDPG with Quantum Price Levels-Based Trading Strategy","date":"2025-01-15","arxiv_id":"2501.08528","n_code_links":0,"syntology":null},{"paper":null,"title":"An Advantage-based Optimization Method for Reinforcement Learning in Large Action Space","date":"2024-12-17","arxiv_id":"2412.12605","n_code_links":0,"syntology":null},{"paper":null,"title":"Broad Critic Deep Actor Reinforcement Learning for Continuous Control","date":"2024-11-24","arxiv_id":"2411.15806","n_code_links":0,"syntology":null},{"paper":null,"title":"Provably Efficient Action-Manipulation Attack Against Continuous Reinforcement Learning","date":"2024-11-20","arxiv_id":"2411.13116","n_code_links":0,"syntology":null},{"paper":null,"title":"Quantum Policy Gradient in Reproducing Kernel Hilbert Space","date":"2024-11-11","arxiv_id":"2411.06650","n_code_links":0,"syntology":null},{"paper":null,"title":"Mitigating Suboptimality of Deterministic Policy Gradients in Complex Q-functions","date":"2024-10-15","arxiv_id":"2410.11833","n_code_links":0,"syntology":null},{"paper":null,"title":"ETGL-DDPG: A Deep Deterministic Policy Gradient Algorithm for Sparse Reward Continuous Control","date":"2024-10-07","arxiv_id":"2410.05225","n_code_links":0,"syntology":null},{"paper":null,"title":"Robust Deep Reinforcement Learning for Volt-VAR Optimization in Active Distribution System under Uncertainty","date":"2024-09-27","arxiv_id":"2409.18937","n_code_links":0,"syntology":null},{"paper":null,"title":"FH-DRL: Exponential-Hyperbolic Frontier Heuristics with DRL for accelerated Exploration in Unknown Environments","date":"2024-07-26","arxiv_id":"2407.18892","n_code_links":0,"syntology":null},{"paper":null,"title":"The Cross-environment Hyperparameter Setting Benchmark for Reinforcement Learning","date":"2024-07-26","arxiv_id":"2407.18840","n_code_links":0,"syntology":null},{"paper":"/paper/reconfigurable-intelligent-surface-aided-21","title":"Reconfigurable Intelligent Surface Aided Vehicular Edge Computing: Joint Phase-shift Optimization and Multi-User Power Allocation","date":"2024-07-18","arxiv_id":"2407.13123","n_code_links":1,"syntology":null},{"paper":null,"title":"Deep reinforcement learning with symmetric data augmentation applied for aircraft lateral attitude tracking control","date":"2024-07-13","arxiv_id":"2407.11077","n_code_links":0,"syntology":null},{"paper":null,"title":"Deep Reinforcement Learning Strategies in Finance: Insights into Asset Holding, Trading Behavior, and Purchase Diversity","date":"2024-06-29","arxiv_id":"2407.09557","n_code_links":0,"syntology":null},{"paper":null,"title":"Performance Comparison of Deep RL Algorithms for Mixed Traffic Cooperative Lane-Changing","date":"2024-06-25","arxiv_id":"2407.02521","n_code_links":0,"syntology":null},{"paper":null,"title":"Towards Dynamic Resource Allocation and Client Scheduling in Hierarchical Federated Learning: A Two-Phase Deep Reinforcement Learning Approach","date":"2024-06-21","arxiv_id":"2406.14910","n_code_links":0,"syntology":null},{"paper":null,"title":"Value Improved Actor Critic Algorithms","date":"2024-06-03","arxiv_id":"2406.01423","n_code_links":0,"syntology":null},{"paper":null,"title":"Meta Reinforcement Learning for Resource Allocation in Multi-Antenna UAV Network with Rate Splitting Multiple Access","date":"2024-05-18","arxiv_id":"2405.11306","n_code_links":0,"syntology":null},{"paper":null,"title":"Continuous Control Reinforcement Learning: Distributed Distributional DrQ Algorithms","date":"2024-04-16","arxiv_id":"2404.10645","n_code_links":0,"syntology":null},{"paper":null,"title":"Conservative DDPG -- Pessimistic RL without Ensemble","date":"2024-03-08","arxiv_id":"2403.05732","n_code_links":0,"syntology":null},{"paper":null,"title":"Fill-and-Spill: Deep Reinforcement Learning Policy Gradient Methods for Reservoir Operation Decision and Control","date":"2024-03-07","arxiv_id":"2403.04195","n_code_links":0,"syntology":null}],"papers_shown":30,"tasks":[{"task":"/task/reinforcement-learning-1","name":"Reinforcement Learning (RL)","papers":131},{"task":"/task/reinforcement-learning-2","name":"reinforcement-learning","papers":114},{"task":"/task/deep-reinforcement-learning","name":"Deep Reinforcement Learning","papers":105},{"task":"/task/reinforcement-learning","name":"Reinforcement Learning","papers":102},{"task":"/task/continuous-control","name":"Continuous Control","papers":36},{"task":"/task/continuous-control","name":"continuous-control","papers":34},{"task":"/task/q-learning","name":"Q-Learning","papers":17},{"task":"/task/mujoco","name":"MuJoCo","papers":16},{"task":"/task/openai-gym","name":"OpenAI Gym","papers":15},{"task":"/task/decision-making","name":"Decision Making","papers":14},{"task":"/task/management","name":"Management","papers":13},{"task":"/task/autonomous-driving","name":"Autonomous Driving","papers":8},{"task":"/task/multi-agent-reinforcement-learning","name":"Multi-agent Reinforcement Learning","papers":6},{"task":"/task/scheduling","name":"Scheduling","papers":6},{"task":"/task/energy-management","name":"energy management","papers":6},{"task":"/task/diversity","name":"Diversity","papers":5},{"task":"/task/federated-learning","name":"Federated Learning","papers":5},{"task":"/task/benchmarking","name":"Benchmarking","papers":4},{"task":"/task/data-augmentation","name":"Data Augmentation","papers":4},{"task":"/task/edge-computing","name":"Edge-computing","papers":4}],"tasks_shown":20,"n_tasks":88,"usage_by_year":[{"year":"2015","papers":1},{"year":"2016","papers":1},{"year":"2017","papers":6},{"year":"2018","papers":21},{"year":"2019","papers":22},{"year":"2020","papers":34},{"year":"2021","papers":31},{"year":"2022","papers":32},{"year":"2023","papers":33},{"year":"2024","papers":26},{"year":"2025","papers":11}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/ddpg"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}