{"url":"/method/soft-actor-critic","slug":"soft-actor-critic","name":"Soft Actor Critic","full_name":"Soft Actor Critic","full_name_withheld":false,"description_markdown":"**Soft Actor Critic**, or **SAC**, is an off-policy actor-critic deep RL algorithm based on the maximum entropy reinforcement learning framework. In this framework, the actor aims to maximize expected reward while also maximizing entropy. That is, to succeed at the task while acting as randomly as possible. Prior deep RL methods based on this framework have been formulated as [Q-learning methods](https://paperswithcode.com/method/q-learning). [SAC](https://paperswithcode.com/method/sac) combines off-policy updates with a stable stochastic actor-critic formulation.\r\n\r\nThe SAC objective has a number of advantages. First, the policy is incentivized to explore more widely, while giving up on clearly unpromising avenues. Second, the policy can capture multiple modes of near-optimal behavior. In problem settings where multiple actions seem equally attractive, the policy will commit equal probability mass to those actions. Lastly, the authors present evidence that it improves learning speed over state-of-art methods that optimize the conventional RL objective function.","description_state":"present","introduced_year":null,"introduced_by":{"title":"Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor","paper":"/paper/soft-actor-critic-off-policy-maximum-entropy","first_author":"Tuomas Haarnoja","n_authors":4,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/soft-actor-critic-off-policy-maximum-entropy"},"source":{"url":"http://arxiv.org/abs/1801.01290v2","title":"Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Reinforcement Learning","area_id":"reinforcement-learning","collection":"Policy Gradient Methods","url":"/methods/category/policy-gradient-methods","pwc_aliases":[]}],"n_papers_tagged":58,"archive_num_papers":58,"papers_newest_first":[{"paper":null,"title":"Moderate Actor-Critic Methods: Controlling Overestimation Bias via Expectile Loss","date":"2025-04-14","arxiv_id":"2504.09929","n_code_links":0,"syntology":null},{"paper":null,"title":"Closing the Intent-to-Behavior Gap via Fulfillment Priority Logic","date":"2025-03-04","arxiv_id":"2503.05818","n_code_links":0,"syntology":null},{"paper":null,"title":"IL-SOAR : Imitation Learning with Soft Optimistic Actor cRitic","date":"2025-02-27","arxiv_id":"2502.19859","n_code_links":0,"syntology":null},{"paper":"/paper/langevin-soft-actor-critic-efficient","title":"Langevin Soft Actor-Critic: Efficient Exploration through Uncertainty-Driven Critic Learning","date":"2025-01-29","arxiv_id":"2501.17827","n_code_links":1,"syntology":{"ran":6,"of":11,"unverified":5,"pointer_only":0}},{"paper":"/paper/reinforcement-learning-controlled-adaptive","title":"Reinforcement Learning Controlled Adaptive PSO for Task Offloading in IIoT Edge Computing","date":"2025-01-25","arxiv_id":"2501.15203","n_code_links":1,"syntology":null},{"paper":null,"title":"Average Reward Reinforcement Learning for Wireless Radio Resource Management","date":"2025-01-12","arxiv_id":"2501.06700","n_code_links":0,"syntology":null},{"paper":null,"title":"Learn 2 Rage: Experiencing The Emotional Roller Coaster That Is Reinforcement Learning","date":"2024-10-24","arxiv_id":"2410.18462","n_code_links":0,"syntology":null},{"paper":null,"title":"Augmented Lagrangian-Based Safe Reinforcement Learning Approach for Distribution System Volt/VAR Control","date":"2024-10-19","arxiv_id":"2410.15188","n_code_links":0,"syntology":null},{"paper":null,"title":"Solving The Dynamic Volatility Fitting Problem: A Deep Reinforcement Learning Approach","date":"2024-10-15","arxiv_id":"2410.11789","n_code_links":0,"syntology":null},{"paper":"/paper/deep-attention-driven-reinforcement-learning","title":"Deep Attention Driven Reinforcement Learning (DAD-RL) for Autonomous Decision-Making in Dynamic Environment","date":"2024-07-12","arxiv_id":"2407.08932","n_code_links":1,"syntology":null},{"paper":null,"title":"Real-time system optimal traffic routing under uncertainties -- Can physics models boost reinforcement learning?","date":"2024-07-10","arxiv_id":"2407.07364","n_code_links":0,"syntology":null},{"paper":null,"title":"Enhanced Safety in Autonomous Driving: Integrating Latent State Diffusion Model for End-to-End Navigation","date":"2024-07-08","arxiv_id":"2407.06317","n_code_links":0,"syntology":null},{"paper":"/paper/a-fast-balance-optimization-approach-for","title":"A fast balance optimization approach for charging enhancement of lithium-ion battery packs through deep reinforcement learning","date":"2024-04-24","arxiv_id":null,"n_code_links":1,"syntology":null},{"paper":null,"title":"Imitation Game: A Model-based and Imitation Learning Deep Reinforcement Learning Hybrid","date":"2024-04-02","arxiv_id":"2404.01794","n_code_links":0,"syntology":null},{"paper":null,"title":"K-percent Evaluation for Lifelong RL","date":"2024-04-02","arxiv_id":"2404.02113","n_code_links":0,"syntology":null},{"paper":null,"title":"Deep Reinforcement Learning for Local Path Following of an Autonomous Formula SAE Vehicle","date":"2024-01-05","arxiv_id":"2401.02903","n_code_links":0,"syntology":null},{"paper":null,"title":"Dynamic Fairness-Aware Spectrum Auction for Enhanced Licensed Shared Access in 6G Networks","date":"2023-12-20","arxiv_id":"2312.12867","n_code_links":0,"syntology":null},{"paper":null,"title":"On Designing Multi-UAV aided Wireless Powered Dynamic Communication via Hierarchical Deep Reinforcement Learning","date":"2023-12-13","arxiv_id":"2312.07917","n_code_links":0,"syntology":null},{"paper":null,"title":"Joint Sensing and Communication Optimization in Target-Mounted STARS-Assisted Vehicular Networks: A MADRL Approach","date":"2023-11-17","arxiv_id":"2311.10352","n_code_links":0,"syntology":null},{"paper":"/paper/belief-projection-based-reinforcement","title":"Belief Projection-Based Reinforcement Learning for Environments with Delayed Feedback","date":"2023-09-21","arxiv_id":null,"n_code_links":1,"syntology":null},{"paper":null,"title":"Hybrid of representation learning and reinforcement learning for dynamic and complex robotic motion planning","date":"2023-09-07","arxiv_id":"2309.03758","n_code_links":0,"syntology":null},{"paper":null,"title":"A Safe Deep Reinforcement Learning Approach for Energy Efficient Federated Learning in Wireless Communication Networks","date":"2023-08-21","arxiv_id":"2308.10664","n_code_links":0,"syntology":null},{"paper":"/paper/hint-assisted-reinforcement-learning-an","title":"Hint assisted reinforcement learning: an application in radio astronomy","date":"2023-01-10","arxiv_id":"2301.03933","n_code_links":1,"syntology":null},{"paper":null,"title":"Efficient Exploration in Resource-Restricted Reinforcement Learning","date":"2022-12-14","arxiv_id":"2212.06988","n_code_links":0,"syntology":null},{"paper":null,"title":"RL-Based Guidance in Outpatient Hysteroscopy Training: A Feasibility Study","date":"2022-11-26","arxiv_id":"2211.14541","n_code_links":0,"syntology":null},{"paper":null,"title":"A Deep Reinforcement Learning-Based Charging Scheduling Approach with Augmented Lagrangian for Electric Vehicle","date":"2022-09-20","arxiv_id":"2209.09772","n_code_links":0,"syntology":null},{"paper":null,"title":"Plug-and-Play Model-Agnostic Counterfactual Policy Synthesis for Deep Reinforcement Learning based Recommendation","date":"2022-08-10","arxiv_id":"2208.05142","n_code_links":0,"syntology":null},{"paper":"/paper/imitation-learning-by-state-only-distribution","title":"Imitation Learning by State-Only Distribution Matching","date":"2022-02-09","arxiv_id":"2202.04332","n_code_links":1,"syntology":{"ran":1,"of":1,"unverified":0,"pointer_only":0}},{"paper":"/paper/smart-magnetic-microrobots-learn-to-swim-with","title":"Smart Magnetic Microrobots Learn to Swim with Deep Reinforcement Learning","date":"2022-01-14","arxiv_id":"2201.05599","n_code_links":1,"syntology":null},{"paper":null,"title":"Actor Loss of Soft Actor Critic Explained","date":"2021-12-31","arxiv_id":"2112.15568","n_code_links":0,"syntology":null}],"papers_shown":30,"tasks":[{"task":"/task/reinforcement-learning-1","name":"Reinforcement Learning (RL)","papers":32},{"task":"/task/reinforcement-learning","name":"Reinforcement Learning","papers":28},{"task":"/task/reinforcement-learning-2","name":"reinforcement-learning","papers":25},{"task":"/task/deep-reinforcement-learning","name":"Deep Reinforcement Learning","papers":22},{"task":"/task/continuous-control","name":"Continuous Control","papers":14},{"task":"/task/continuous-control","name":"continuous-control","papers":12},{"task":"/task/decision-making","name":"Decision Making","papers":6},{"task":"/task/imitation-learning","name":"Imitation Learning","papers":5},{"task":"/task/openai-gym","name":"OpenAI Gym","papers":5},{"task":"/task/efficient-exploration","name":"Efficient Exploration","papers":4},{"task":"/task/q-learning","name":"Q-Learning","papers":4},{"task":"/task/motion-planning","name":"Motion Planning","papers":3},{"task":"/task/autonomous-driving","name":"Autonomous Driving","papers":2},{"task":"/task/autonomous-racing","name":"Autonomous Racing","papers":2},{"task":"/task/data-augmentation","name":"Data Augmentation","papers":2},{"task":"/task/management","name":"Management","papers":2},{"task":"/task/model-based-reinforcement-learning","name":"Model-based Reinforcement Learning","papers":2},{"task":"/task/mujoco","name":"MuJoCo","papers":2},{"task":"/task/astronomy","name":"Astronomy","papers":1},{"task":"/task/atari-games","name":"Atari Games","papers":1}],"tasks_shown":20,"n_tasks":55,"usage_by_year":[{"year":"2018","papers":2},{"year":"2019","papers":8},{"year":"2020","papers":11},{"year":"2021","papers":8},{"year":"2022","papers":6},{"year":"2023","papers":7},{"year":"2024","papers":10},{"year":"2025","papers":6}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/soft-actor-critic"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}