Methods › Reinforcement Learning › Policy Gradient Methods › D4PG
Distributed Distributional DDPG
D4PG
Introduced by Gabriel Barth-Maron et al. in Distributed Distributional Deterministic Policy Gradients
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
D4PG, or Distributed Distributional DDPG, is a policy gradient algorithm that extends upon the DDPG. The improvements include a distributional updates to the DDPG algorithm, combined with the use of multiple distributed workers all writing into the same replay table. The biggest performance gain of other simpler changes was the use of N-step returns. The authors found that the use of prioritized experience replay was less crucial to the overall D4PG algorithm especially on harder problems.
Papers archive 2025-07-28
11 shown of 11, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
Learning in complex action spaces without policy gradients 8 Oct 2024 · 0 repositories · arXiv:2410.06317
-
Mitigating Estimation Errors by Twin TD-Regularized Actor and Critic for Deep Reinforcement Learning 7 Nov 2023 · 0 repositories · arXiv:2311.03711
-
SDGym: Low-Code Reinforcement Learning Environments using System Dynamics Models 19 Oct 2023 · 1 repository · arXiv:2310.12494
-
A Long N-step Surrogate Stage Reward for Deep Reinforcement Learning 21 Sep 2023 · 0 repositories
-
Gamma and Vega Hedging Using Deep Distributional Reinforcement Learning 10 May 2022 · 1 repository · arXiv:2205.05614
-
Revisiting Gaussian mixture critics in off-policy reinforcement learning: a sample-based approach 21 Apr 2022 · 1 repository · arXiv:2204.10256
-
Tonic: A Deep Reinforcement Learning Library for Fast Prototyping and Benchmarking 15 Nov 2020 · 1 repository · arXiv:2011.07537Syntology ran 3 of 4 samples · 1 unverified
-
Distributed Uplink Beamforming in Cell-Free Networks Using Deep Reinforcement Learning 26 Jun 2020 · 0 repositories · arXiv:2006.15138
-
Sample-based Distributional Policy Gradient 8 Jan 2020 · 0 repositories · arXiv:2001.02652
-
TF-Replicator: Distributed Machine Learning for Researchers 1 Feb 2019 · 1 repository · arXiv:1902.00465Syntology ran 2 of 3 samples · 1 unverified
-
Distributed Distributional Deterministic Policy Gradients 23 Apr 2018 · 5 repositories · arXiv:1804.08617
Tasks archive 2025-07-28
15 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections