Methods › General › Regularization › Entropy Regularization
Entropy Regularization
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
Entropy Regularization is a type of regularization used in reinforcement learning. For on-policy policy gradient based methods like A3C, the same mutual reinforcement behaviour leads to a highly-peaked π(a|s) towards a few actions or action sequences, since it is easier for the actor and critic to overoptimise to a small portion of the environment. To reduce this problem, entropy regularization adds an entropy term to the loss to promote action diversity:
H(X) = -∑π(x)log(π(x))
Image Credit: Wikipedia
Papers archive 2025-07-28
30 shown of 1,128, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
Viability of Future Actions: Robust Safety in Reinforcement Learning via Entropy Regularization 12 Jun 2025 · 1 repository · arXiv:2506.10871
-
Ego-centric Learning of Communicative World Models for Autonomous Driving 9 Jun 2025 · 0 repositories · arXiv:2506.08149
-
Autonomous Vehicle Lateral Control Using Deep Reinforcement Learning with MPC-PID Demonstration 4 Jun 2025 · 0 repositories · arXiv:2506.04040
-
PPO in the Fisher-Rao geometry 4 Jun 2025 · 0 repositories · arXiv:2506.03757
-
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming 4 Jun 2025 · 1 repository · arXiv:2506.04302
-
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective 3 Jun 2025 · 0 repositories · arXiv:2506.02553
-
Rewarding the Unlikely: Lifting GRPO Beyond Distribution Sharpening 3 Jun 2025 · 0 repositories · arXiv:2506.02355
-
The Surprising Effectiveness of Negative Reinforcement in LLM Reasoning 2 Jun 2025 · 1 repository · arXiv:2506.01347Syntology ran 3 of 3 samples · 0 unverified
-
DriveMind: A Dual-VLM based Reinforcement Learning Framework for Autonomous Driving 1 Jun 2025 · 0 repositories · arXiv:2506.00819
-
Language-Guided Multi-Agent Learning in Simulations: A Unified Framework and Evaluation 1 Jun 2025 · 0 repositories · arXiv:2506.04251
-
Using Diffusion Ensembles to Estimate Uncertainty for End-to-End Autonomous Driving 31 May 2025 · 0 repositories · arXiv:2506.00560
-
AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning 30 May 2025 · 1 repository · arXiv:2505.24298Syntology ran 0 of 13 samples · 13 unverified
-
Composite Reward Design in PPO-Driven Adaptive Filtering 29 May 2025 · 1 repository · arXiv:2506.06323
-
Segment Policy Optimization: Effective Segment-Level Credit Assignment in RL for Large Language Models 29 May 2025 · 1 repository · arXiv:2505.23564Syntology ran 7 of 9 samples · 2 unverified
-
TSENOR: Highly-Efficient Algorithm for Finding Transposable N:M Sparse Masks 29 May 2025 · 0 repositories · arXiv:2505.23949
-
R1-Code-Interpreter: Training LLMs to Reason with Code via Supervised and Reinforcement Learning 27 May 2025 · 1 repository · arXiv:2505.21668
-
What Can RL Bring to VLA Generalization? An Empirical Study 26 May 2025 · 0 repositories · arXiv:2505.19789
-
Improving Value Estimation Critically Enhances Vanilla Policy Gradient 25 May 2025 · 1 repository · arXiv:2505.19247
-
GenPO: Generative Diffusion Models Meet On-Policy Reinforcement Learning 24 May 2025 · 0 repositories · arXiv:2505.18763
-
A Robust PPO-optimized Tabular Transformer Framework for Intrusion Detection in Industrial IoT Systems 23 May 2025 · 1 repository · arXiv:2505.18234
-
ShIOEnv: A CLI Behavior-Capturing Environment Enabling Grammar-Guided Command Synthesis for Dataset Curation 23 May 2025 · 1 repository · arXiv:2505.18374
-
The Cell Must Go On: Agar.io for Continual Reinforcement Learning 23 May 2025 · 1 repository · arXiv:2505.18347
-
Raw2Drive: Reinforcement Learning with Aligned World Models for End-to-End Autonomous Driving (in CARLA v2) 22 May 2025 · 0 repositories · arXiv:2505.16394
-
AM-PPO: (Advantage) Alpha-Modulation with Proximal Policy Optimization 21 May 2025 · 0 repositories · arXiv:2505.15514
-
HCRMP: A LLM-Hinted Contextual Reinforcement Learning Framework for Autonomous Driving 21 May 2025 · 0 repositories · arXiv:2505.15793
-
iPad: Iterative Proposal-centric End-to-End Autonomous Driving 21 May 2025 · 1 repository · arXiv:2505.15111
-
Trajectory Bellman Residual Minimization: A Simple Value-Based Method for LLM Reasoning 21 May 2025 · 0 repositories · arXiv:2505.15311Syntology ran 2 of 2 samples · 0 unverified
-
Interpretable Reinforcement Learning for Load Balancing using Kolmogorov-Arnold Networks 20 May 2025 · 0 repositories · arXiv:2505.14459
-
KIPPO: Koopman-Inspired Proximal Policy Optimization 20 May 2025 · 0 repositories · arXiv:2505.14566
-
PEER pressure: Model-to-Model Regularization for Single Source Domain Generalization 19 May 2025 · 0 repositories · arXiv:2505.12745
Tasks archive 2025-07-28
20 shown of 404 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections