Methods › Reinforcement Learning › Policy Gradient Methods › DPG
Deterministic Policy Gradient
DPG
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
Deterministic Policy Gradient, or DPG, is a policy gradient method for reinforcement learning. Instead of the policy function π(.|s) being modeled as a probability distribution, DPG considers and calculates gradients for a deterministic policy a = μₜₕₑₜₐ(s).
Papers archive 2025-07-28
20 shown of 20, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
DPG loss functions for learning parameter-to-solution maps by neural networks 23 Jun 2025 · 0 repositories · arXiv:2506.18773
-
ReasonGen-R1: CoT for Autoregressive Image generation models through SFT and RL 30 May 2025 · 1 repository · arXiv:2505.24875Syntology ran 5 of 13 samples · 8 unverified
-
Harnessing Caption Detailness for Data-Efficient Text-to-Image Generation 21 May 2025 · 0 repositories · arXiv:2505.15172
-
DIP-Watermark: A Double Identity Protection Method Based on Robust Adversarial Watermark 23 Apr 2024 · 0 repositories · arXiv:2404.14693
-
Dynamic Generation of Personalities with Large Language Models 10 Apr 2024 · 1 repository · arXiv:2404.07084
-
Decision Predicate Graphs: Enhancing Interpretability in Tree Ensembles 3 Apr 2024 · 0 repositories · arXiv:2404.02942
-
BRAIn: Bayesian Reward-conditioned Amortized Inference for natural language generation from feedback 4 Feb 2024 · 0 repositories · arXiv:2402.02479
-
Dual-based Online Learning of Dynamic Network Topologies 14 Nov 2022 · 0 repositories · arXiv:2211.07449
-
Dynamic Sparse R-CNN 4 May 2022 · 0 repositories · arXiv:2205.02101
-
Controlling Conditional Language Models without Catastrophic Forgetting 1 Dec 2021 · 2 repositories · arXiv:2112.00791Syntology ran 3 of 8 samples · 5 unverified · 4 pointer-only (licence)
-
MPC-based Reinforcement Learning for a Simplified Freight Mission of Autonomous Surface Vehicles 16 Jun 2021 · 0 repositories · arXiv:2106.08634
-
Deep Reinforcement Agent for Scheduling in HPC 11 Feb 2021 · 1 repository · arXiv:2102.06243
-
A review of motion planning algorithms for intelligent robotics 4 Feb 2021 · 0 repositories · arXiv:2102.02376
-
OffCon³: What is state of the art anyway? 27 Jan 2021 · 1 repository · arXiv:2101.11331
-
Zeroth-order Deterministic Policy Gradient 12 Jun 2020 · 0 repositories · arXiv:2006.07314
-
Investigation on the generalization of the Sampled Policy Gradient algorithm 9 Oct 2019 · 0 repositories · arXiv:1910.03728
-
Sampled Policy Gradient for Learning to Play the Game Agar.io 15 Sep 2018 · 2 repositories · arXiv:1809.05763
-
Directed Policy Gradient for Safe Reinforcement Learning with Human Advice 13 Aug 2018 · 0 repositories · arXiv:1808.04096
-
Deterministic Policy Gradients With General State Transitions 10 Jul 2018 · 0 repositories · arXiv:1807.03708
-
Expected Policy Gradients 15 Jun 2017 · 0 repositories · arXiv:1706.05374
Tasks archive 2025-07-28
20 shown of 28 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections