Papers › Particle Based Stochastic Policy Optimization
Particle Based Stochastic Policy Optimization
Qiwei Ye, Yuxuan Song, Chang Liu, Fangyun Wei, Tao Qin, Tie-Yan Liu
Stochastic polic have been widely applied for their good property in exploration and uncertainty quantification. Modeling policy distribution by joint state-action distribution within the exponential family has enabled flexibility in exploration and learning multi-modal policies and also involved the probabilistic perspective of deep reinforcement learning (RL). The connection between probabilistic inference and RL makes it possible to leverage the advancements of probabilistic optimization tools. However, recent efforts are limited to the minimization of reverse KLdivergence which is confidence-seeking and may fade the merit of a stochastic policy. To leverage the full potential of stochastic policy and provide more flexible property, there is a strong motivation to consider different update rules during policy optimization. In this paper, we propose a particle-based probabilistic pol-icy optimization framework, ParPI, which enables the usage of a broad family of divergence or distances, such asf-divergences, and the Wasserstein distance which could serve better probabilistic behavior of the learned stochastic policy. Experiments in both online and offline settings demonstrate the effectiveness of the proposed algorithm as well as the characteristics of different discrepancy measures for policy optimization.
Code
No code repository is listed for this paper in the archive or in Syntology's graph.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| MuJoCo Games | Ant-v3 | ParPI | Average Reward | 5142 | #1 of 1 | Archive leaderboard | report |
| MuJoCo Games | HalfCHeetah-v3 | ParPI | Average Reward | 11738 | #1 of 1 | Archive leaderboard | report |
| MuJoCo Games | Hopper-v3 | ParPI | Average Reward | 3042 | #1 of 1 | Archive leaderboard | report |
| MuJoCo Games | Humanoid-v3 | ParPI | Average Reward | 4912 | #1 of 1 | Archive leaderboard | report |
| MuJoCo Games | Walker2d-v3 | ParPI | Average Reward | 5201 | #1 of 1 | Archive leaderboard | report |
| Offline RL | Walker2d | ParPI | D4RL Normalized Score | 151.4 | #1 of 1 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections