Methods › Reinforcement Learning › Policy Gradient Methods › Soft Actor Critic
Soft Actor Critic
Introduced by Tuomas Haarnoja et al. in Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
Soft Actor Critic, or SAC, is an off-policy actor-critic deep RL algorithm based on the maximum entropy reinforcement learning framework. In this framework, the actor aims to maximize expected reward while also maximizing entropy. That is, to succeed at the task while acting as randomly as possible. Prior deep RL methods based on this framework have been formulated as Q-learning methods. SAC combines off-policy updates with a stable stochastic actor-critic formulation.
The SAC objective has a number of advantages. First, the policy is incentivized to explore more widely, while giving up on clearly unpromising avenues. Second, the policy can capture multiple modes of near-optimal behavior. In problem settings where multiple actions seem equally attractive, the policy will commit equal probability mass to those actions. Lastly, the authors present evidence that it improves learning speed over state-of-art methods that optimize the conventional RL objective function.
Papers archive 2025-07-28
30 shown of 58, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
Moderate Actor-Critic Methods: Controlling Overestimation Bias via Expectile Loss 14 Apr 2025 · 0 repositories · arXiv:2504.09929
-
Closing the Intent-to-Behavior Gap via Fulfillment Priority Logic 4 Mar 2025 · 0 repositories · arXiv:2503.05818
-
IL-SOAR : Imitation Learning with Soft Optimistic Actor cRitic 27 Feb 2025 · 0 repositories · arXiv:2502.19859
-
Langevin Soft Actor-Critic: Efficient Exploration through Uncertainty-Driven Critic Learning 29 Jan 2025 · 1 repository · arXiv:2501.17827Syntology ran 6 of 11 samples · 5 unverified
-
Reinforcement Learning Controlled Adaptive PSO for Task Offloading in IIoT Edge Computing 25 Jan 2025 · 1 repository · arXiv:2501.15203
-
Average Reward Reinforcement Learning for Wireless Radio Resource Management 12 Jan 2025 · 0 repositories · arXiv:2501.06700
-
Learn 2 Rage: Experiencing The Emotional Roller Coaster That Is Reinforcement Learning 24 Oct 2024 · 0 repositories · arXiv:2410.18462
-
Augmented Lagrangian-Based Safe Reinforcement Learning Approach for Distribution System Volt/VAR Control 19 Oct 2024 · 0 repositories · arXiv:2410.15188
-
Solving The Dynamic Volatility Fitting Problem: A Deep Reinforcement Learning Approach 15 Oct 2024 · 0 repositories · arXiv:2410.11789
-
Deep Attention Driven Reinforcement Learning (DAD-RL) for Autonomous Decision-Making in Dynamic Environment 12 Jul 2024 · 1 repository · arXiv:2407.08932
-
Real-time system optimal traffic routing under uncertainties -- Can physics models boost reinforcement learning? 10 Jul 2024 · 0 repositories · arXiv:2407.07364
-
Enhanced Safety in Autonomous Driving: Integrating Latent State Diffusion Model for End-to-End Navigation 8 Jul 2024 · 0 repositories · arXiv:2407.06317
-
A fast balance optimization approach for charging enhancement of lithium-ion battery packs through deep reinforcement learning 24 Apr 2024 · 1 repository
-
Imitation Game: A Model-based and Imitation Learning Deep Reinforcement Learning Hybrid 2 Apr 2024 · 0 repositories · arXiv:2404.01794
-
K-percent Evaluation for Lifelong RL 2 Apr 2024 · 0 repositories · arXiv:2404.02113
-
Deep Reinforcement Learning for Local Path Following of an Autonomous Formula SAE Vehicle 5 Jan 2024 · 0 repositories · arXiv:2401.02903
-
Dynamic Fairness-Aware Spectrum Auction for Enhanced Licensed Shared Access in 6G Networks 20 Dec 2023 · 0 repositories · arXiv:2312.12867
-
On Designing Multi-UAV aided Wireless Powered Dynamic Communication via Hierarchical Deep Reinforcement Learning 13 Dec 2023 · 0 repositories · arXiv:2312.07917
-
Joint Sensing and Communication Optimization in Target-Mounted STARS-Assisted Vehicular Networks: A MADRL Approach 17 Nov 2023 · 0 repositories · arXiv:2311.10352
-
Belief Projection-Based Reinforcement Learning for Environments with Delayed Feedback 21 Sep 2023 · 1 repository
-
Hybrid of representation learning and reinforcement learning for dynamic and complex robotic motion planning 7 Sep 2023 · 0 repositories · arXiv:2309.03758
-
A Safe Deep Reinforcement Learning Approach for Energy Efficient Federated Learning in Wireless Communication Networks 21 Aug 2023 · 0 repositories · arXiv:2308.10664
-
Hint assisted reinforcement learning: an application in radio astronomy 10 Jan 2023 · 1 repository · arXiv:2301.03933
-
Efficient Exploration in Resource-Restricted Reinforcement Learning 14 Dec 2022 · 0 repositories · arXiv:2212.06988
-
RL-Based Guidance in Outpatient Hysteroscopy Training: A Feasibility Study 26 Nov 2022 · 0 repositories · arXiv:2211.14541
-
A Deep Reinforcement Learning-Based Charging Scheduling Approach with Augmented Lagrangian for Electric Vehicle 20 Sep 2022 · 0 repositories · arXiv:2209.09772
-
Plug-and-Play Model-Agnostic Counterfactual Policy Synthesis for Deep Reinforcement Learning based Recommendation 10 Aug 2022 · 0 repositories · arXiv:2208.05142
-
Imitation Learning by State-Only Distribution Matching 9 Feb 2022 · 1 repository · arXiv:2202.04332Syntology ran 1 of 1 samples · 0 unverified
-
Smart Magnetic Microrobots Learn to Swim with Deep Reinforcement Learning 14 Jan 2022 · 1 repository · arXiv:2201.05599
-
Actor Loss of Soft Actor Critic Explained 31 Dec 2021 · 0 repositories · arXiv:2112.15568
Tasks archive 2025-07-28
20 shown of 55 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections