{"url":"/method/d4pg","slug":"d4pg","name":"D4PG","full_name":"Distributed Distributional DDPG","full_name_withheld":false,"description_markdown":"**D4PG**, or **Distributed Distributional DDPG**, is a policy gradient algorithm that extends upon the [DDPG](https://paperswithcode.com/method/ddpg). The improvements include a distributional updates to the DDPG algorithm, combined with the use of multiple distributed workers all writing into the same replay table. The biggest performance gain of other simpler changes was the use of $N$-step returns. The authors found that the use of [prioritized experience replay](https://paperswithcode.com/method/prioritized-experience-replay) was less crucial to the overall D4PG algorithm especially on harder problems.","description_state":"present","introduced_year":null,"introduced_by":{"title":"Distributed Distributional Deterministic Policy Gradients","paper":"/paper/distributed-distributional-deterministic","first_author":"Gabriel Barth-Maron","n_authors":9,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/distributed-distributional-deterministic"},"source":{"url":"http://arxiv.org/abs/1804.08617v1","title":"Distributed Distributional Deterministic Policy Gradients","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Reinforcement Learning","area_id":"reinforcement-learning","collection":"Policy Gradient Methods","url":"/methods/category/policy-gradient-methods","pwc_aliases":[]}],"n_papers_tagged":11,"archive_num_papers":11,"papers_newest_first":[{"paper":null,"title":"Learning in complex action spaces without policy gradients","date":"2024-10-08","arxiv_id":"2410.06317","n_code_links":0,"syntology":null},{"paper":null,"title":"Mitigating Estimation Errors by Twin TD-Regularized Actor and Critic for Deep Reinforcement Learning","date":"2023-11-07","arxiv_id":"2311.03711","n_code_links":0,"syntology":null},{"paper":"/paper/sdgym-low-code-reinforcement-learning","title":"SDGym: Low-Code Reinforcement Learning Environments using System Dynamics Models","date":"2023-10-19","arxiv_id":"2310.12494","n_code_links":1,"syntology":null},{"paper":null,"title":"A Long $N$-step Surrogate Stage Reward for Deep Reinforcement Learning","date":"2023-09-21","arxiv_id":null,"n_code_links":0,"syntology":null},{"paper":"/paper/gamma-and-vega-hedging-using-deep","title":"Gamma and Vega Hedging Using Deep Distributional Reinforcement Learning","date":"2022-05-10","arxiv_id":"2205.05614","n_code_links":1,"syntology":null},{"paper":"/paper/revisiting-gaussian-mixture-critic-in-off","title":"Revisiting Gaussian mixture critics in off-policy reinforcement learning: a sample-based approach","date":"2022-04-21","arxiv_id":"2204.10256","n_code_links":1,"syntology":null},{"paper":"/paper/tonic-a-deep-reinforcement-learning-library","title":"Tonic: A Deep Reinforcement Learning Library for Fast Prototyping and Benchmarking","date":"2020-11-15","arxiv_id":"2011.07537","n_code_links":1,"syntology":{"ran":3,"of":4,"unverified":1,"pointer_only":0}},{"paper":null,"title":"Distributed Uplink Beamforming in Cell-Free Networks Using Deep Reinforcement Learning","date":"2020-06-26","arxiv_id":"2006.15138","n_code_links":0,"syntology":null},{"paper":null,"title":"Sample-based Distributional Policy Gradient","date":"2020-01-08","arxiv_id":"2001.02652","n_code_links":0,"syntology":null},{"paper":"/paper/tf-replicator-distributed-machine-learning","title":"TF-Replicator: Distributed Machine Learning for Researchers","date":"2019-02-01","arxiv_id":"1902.00465","n_code_links":1,"syntology":{"ran":2,"of":3,"unverified":1,"pointer_only":0}},{"paper":"/paper/distributed-distributional-deterministic","title":"Distributed Distributional Deterministic Policy Gradients","date":"2018-04-23","arxiv_id":"1804.08617","n_code_links":5,"syntology":null}],"papers_shown":11,"tasks":[{"task":"/task/reinforcement-learning-1","name":"Reinforcement Learning (RL)","papers":7},{"task":"/task/reinforcement-learning-2","name":"reinforcement-learning","papers":6},{"task":"/task/continuous-control","name":"Continuous Control","papers":4},{"task":"/task/reinforcement-learning","name":"Reinforcement Learning","papers":4},{"task":"/task/continuous-control","name":"continuous-control","papers":4},{"task":"/task/deep-reinforcement-learning","name":"Deep Reinforcement Learning","papers":3},{"task":"/task/openai-gym","name":"OpenAI Gym","papers":3},{"task":"/task/distributional-reinforcement-learning","name":"Distributional Reinforcement Learning","papers":2},{"task":"/task/machine-learning","name":"BIG-bench Machine Learning","papers":1},{"task":"/task/benchmarking","name":"Benchmarking","papers":1},{"task":"/task/image-generation","name":"Image Generation","papers":1},{"task":"/task/policy-gradient-methods","name":"Policy Gradient Methods","papers":1},{"task":null,"name":"Position","papers":1},{"task":"/task/q-learning","name":"Q-Learning","papers":1},{"task":"/task/quantile-regression","name":"quantile regression","papers":1}],"tasks_shown":15,"n_tasks":15,"usage_by_year":[{"year":"2018","papers":1},{"year":"2019","papers":1},{"year":"2020","papers":3},{"year":"2022","papers":2},{"year":"2023","papers":3},{"year":"2024","papers":1}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/d4pg"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}