Methods › Reinforcement Learning › Policy Gradient Methods

Policy Gradient Methods

23 methods 1,633 papers tagged archive 2025-07-28

Policy Gradient Methods try to optimize the policy function directly in reinforcement learning. This contrasts with, for example, Q-Learning, where the policy manifests itself as maximizing a value function. Below you can find a continuously updating catalog of policy gradient methods.

Methods

All 23 methods in this collection, most-tagged first. Year is the archive's introduced_year; the archive stores 2000 when it has none, shown here as “–”. Papers counts distinct papers the archive tags with the method. Click a heading to sort.

PPO Proximal Policy Optimization – 949
DDPG Deep Deterministic Policy Gradient – 218
REINFORCE 1999 185
TD3 Twin Delayed Deep Deterministic – 117
A2C – 82
TRPO Trust Region Policy Optimization – 81
Soft Actor Critic – 58
A3C – 57
MADDPG – 36
DPG Deterministic Policy Gradient 2014 20
IMPALA – 16
ACER – 12
D4PG Distributed Distributional DDPG – 11
Soft Actor-Critic (Autotuned Temperature) – 6
MDPO Mirror Descent Policy Optimization – 4
ACTKR – 2
SVPG Stein Variational Policy Gradient – 2
myGym MyGym: Modular Toolkit for Visuomotor Robotic Tasks – 2
Ape-X DPG – 1
Fisher-BRC – 1
NoisyNet-A3C – 1
Robust Predictable Control – 1
TayPO Taylor Expansion Policy Optimization – 1