Methods › Reinforcement Learning › Policy Gradient Methods
Policy Gradient Methods
Policy Gradient Methods try to optimize the policy function directly in reinforcement learning. This contrasts with, for example, Q-Learning, where the policy manifests itself as maximizing a value function. Below you can find a continuously updating catalog of policy gradient methods.
Methods
All 23 methods in this collection, most-tagged first. Year is the archive's introduced_year; the archive stores 2000 when it has none, shown here as “–”. Papers counts distinct papers the archive tags with the method. Click a heading to sort.
| PPO Proximal Policy Optimization | – | 949 |
| DDPG Deep Deterministic Policy Gradient | – | 218 |
| REINFORCE | 1999 | 185 |
| TD3 Twin Delayed Deep Deterministic | – | 117 |
| A2C | – | 82 |
| TRPO Trust Region Policy Optimization | – | 81 |
| Soft Actor Critic | – | 58 |
| A3C | – | 57 |
| MADDPG | – | 36 |
| DPG Deterministic Policy Gradient | 2014 | 20 |
| IMPALA | – | 16 |
| ACER | – | 12 |
| D4PG Distributed Distributional DDPG | – | 11 |
| Soft Actor-Critic (Autotuned Temperature) | – | 6 |
| MDPO Mirror Descent Policy Optimization | – | 4 |
| ACTKR | – | 2 |
| SVPG Stein Variational Policy Gradient | – | 2 |
| myGym MyGym: Modular Toolkit for Visuomotor Robotic Tasks | – | 2 |
| Ape-X DPG | – | 1 |
| Fisher-BRC | – | 1 |
| NoisyNet-A3C | – | 1 |
| Robust Predictable Control | – | 1 |
| TayPO Taylor Expansion Policy Optimization | – | 1 |