{"url":"/method/a2c","slug":"a2c","name":"A2C","full_name":"A2C","full_name_withheld":false,"description_markdown":"**A2C**, or **Advantage Actor Critic**, is a synchronous version of the [A3C](https://paperswithcode.com/method/a3c) policy gradient method. As an alternative to the asynchronous implementation of A3C, A2C is a synchronous, deterministic implementation that waits for each actor to finish its segment of experience before updating, averaging over all of the actors. This more effectively uses GPUs due to larger batch sizes.\r\n\r\nImage Credit: [OpenAI Baselines](https://openai.com/blog/baselines-acktr-a2c/)","description_state":"present","introduced_year":null,"introduced_by":{"title":"Asynchronous Methods for Deep Reinforcement Learning","paper":"/paper/asynchronous-methods-for-deep-reinforcement","first_author":"Volodymyr Mnih","n_authors":8,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/asynchronous-methods-for-deep-reinforcement"},"source":{"url":"http://arxiv.org/abs/1602.01783v2","title":"Asynchronous Methods for Deep Reinforcement Learning","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Reinforcement Learning","area_id":"reinforcement-learning","collection":"Policy Gradient Methods","url":"/methods/category/policy-gradient-methods","pwc_aliases":[]}],"n_papers_tagged":82,"archive_num_papers":82,"papers_newest_first":[{"paper":null,"title":"Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control","date":"2025-05-13","arxiv_id":"2505.09029","n_code_links":0,"syntology":null},{"paper":null,"title":"Ensemble RL through Classifier Models: Enhancing Risk-Return Trade-offs in Trading Strategies","date":"2025-02-23","arxiv_id":"2502.17518","n_code_links":0,"syntology":null},{"paper":null,"title":"Continuous Learning Conversational AI: A Personalized Agent Framework via A2C Reinforcement Learning","date":"2025-02-18","arxiv_id":"2502.12876","n_code_links":0,"syntology":null},{"paper":"/paper/evorl-a-gpu-accelerated-framework-for","title":"EvoRL: A GPU-accelerated Framework for Evolutionary Reinforcement Learning","date":"2025-01-25","arxiv_id":"2501.15129","n_code_links":2,"syntology":null},{"paper":null,"title":"Innate-Values-driven Reinforcement Learning based Cognitive Modeling","date":"2024-11-14","arxiv_id":"2411.09160","n_code_links":0,"syntology":null},{"paper":null,"title":"Deep Reinforcement Learning Strategies in Finance: Insights into Asset Holding, Trading Behavior, and Purchase Diversity","date":"2024-06-29","arxiv_id":"2407.09557","n_code_links":0,"syntology":null},{"paper":null,"title":"Multistep Criticality Search and Power Shaping in Microreactors with Reinforcement Learning","date":"2024-06-22","arxiv_id":"2406.15931","n_code_links":0,"syntology":null},{"paper":null,"title":"Biological Neurons Compete with Deep Reinforcement Learning in Sample Efficiency in a Simulated Gameworld","date":"2024-05-27","arxiv_id":"2405.16946","n_code_links":0,"syntology":null},{"paper":"/paper/symmetric-reinforcement-learning-loss-for","title":"Symmetric Reinforcement Learning Loss for Robust Learning on Diverse Tasks and Model Scales","date":"2024-05-27","arxiv_id":"2405.17618","n_code_links":1,"syntology":null},{"paper":null,"title":"Extracting Heuristics from Large Language Models for Reward Shaping in Reinforcement Learning","date":"2024-05-24","arxiv_id":"2405.15194","n_code_links":0,"syntology":null},{"paper":null,"title":"Portfolio Management using Deep Reinforcement Learning","date":"2024-05-01","arxiv_id":"2405.01604","n_code_links":0,"syntology":null},{"paper":"/paper/breaching-the-bottleneck-evolutionary","title":"Breaching the Bottleneck: Evolutionary Transition from Reward-Driven Learning to Reward-Agnostic Domain-Adapted Learning in Neuromodulated Neural Nets","date":"2024-04-19","arxiv_id":"2404.12631","n_code_links":1,"syntology":null},{"paper":null,"title":"A2C: A Modular Multi-stage Collaborative Decision Framework for Human-AI Teams","date":"2024-01-25","arxiv_id":"2401.14432","n_code_links":0,"syntology":null},{"paper":null,"title":"Epidemic Decision-making System Based Federated Reinforcement Learning","date":"2023-11-03","arxiv_id":"2311.01749","n_code_links":0,"syntology":null},{"paper":null,"title":"Diagnosis-oriented Medical Image Compression with Efficient Transfer Learning","date":"2023-10-20","arxiv_id":"2310.13250","n_code_links":0,"syntology":null},{"paper":"/paper/deep-reinforcement-learning-based-intelligent-2","title":"Deep Reinforcement Learning-based Intelligent Traffic Signal Controls with Optimized CO2 emissions","date":"2023-10-19","arxiv_id":"2310.13129","n_code_links":1,"syntology":null},{"paper":null,"title":"Solving the Quadratic Assignment Problem using Deep Reinforcement Learning","date":"2023-10-02","arxiv_id":"2310.01604","n_code_links":0,"syntology":null},{"paper":null,"title":"Raijū: Reinforcement Learning-Guided Post-Exploitation for Automating Security Assessment of Network Systems","date":"2023-09-27","arxiv_id":"2309.15518","n_code_links":0,"syntology":null},{"paper":null,"title":"SAF-Net: Self-Attention Fusion Network for Myocardial Infarction Detection using Multi-View Echocardiography","date":"2023-09-27","arxiv_id":"2309.15520","n_code_links":0,"syntology":null},{"paper":null,"title":"Career Path Recommendations for Long-term Income Maximization: A Reinforcement Learning Approach","date":"2023-09-11","arxiv_id":"2309.05391","n_code_links":0,"syntology":null},{"paper":null,"title":"Semantic Consistency for Assuring Reliability of Large Language Models","date":"2023-08-17","arxiv_id":"2308.09138","n_code_links":0,"syntology":null},{"paper":null,"title":"Deep Reinforcement Learning for ESG financial portfolio management","date":"2023-06-19","arxiv_id":"2307.09631","n_code_links":0,"syntology":null},{"paper":null,"title":"Multi-Agent Reinforcement Learning for Network Routing in Integrated Access Backhaul Networks","date":"2023-05-12","arxiv_id":"2305.16170","n_code_links":0,"syntology":null},{"paper":null,"title":"Deep reinforcement learning applied to an assembly sequence planning problem with user preferences","date":"2023-04-13","arxiv_id":"2304.06567","n_code_links":0,"syntology":null},{"paper":null,"title":"Read and Reap the Rewards: Learning to Play Atari with the Help of Instruction Manuals","date":"2023-02-09","arxiv_id":"2302.04449","n_code_links":0,"syntology":null},{"paper":null,"title":"A Scale-Independent Multi-Objective Reinforcement Learning with Convergence Analysis","date":"2023-02-08","arxiv_id":"2302.04179","n_code_links":0,"syntology":null},{"paper":"/paper/learning-fast-and-slow-a-goal-directed-memory","title":"Learning, Fast and Slow: A Goal-Directed Memory-Based Approach for Dynamic Environments","date":"2023-01-31","arxiv_id":"2301.13758","n_code_links":1,"syntology":null},{"paper":null,"title":"Towards automating Codenames spymasters with deep reinforcement learning","date":"2022-12-28","arxiv_id":"2212.14104","n_code_links":0,"syntology":null},{"paper":"/paper/solving-the-side-chain-packing-arrangement-of","title":"Reinforcement Learning for Molecular Dynamics Optimization: A Stochastic Pontryagin Maximum Principle Approach","date":"2022-12-06","arxiv_id":"2212.03320","n_code_links":1,"syntology":null},{"paper":null,"title":"Towards More Efficient Shared Autonomous Mobility: A Learning-Based Fleet Repositioning Approach","date":"2022-10-16","arxiv_id":"2210.08659","n_code_links":0,"syntology":null}],"papers_shown":30,"tasks":[{"task":"/task/reinforcement-learning-1","name":"Reinforcement Learning (RL)","papers":52},{"task":"/task/reinforcement-learning-2","name":"reinforcement-learning","papers":50},{"task":"/task/reinforcement-learning","name":"Reinforcement Learning","papers":49},{"task":"/task/deep-reinforcement-learning","name":"Deep Reinforcement Learning","papers":30},{"task":"/task/decision-making","name":"Decision Making","papers":12},{"task":"/task/atari-games","name":"Atari Games","papers":10},{"task":"/task/q-learning","name":"Q-Learning","papers":7},{"task":"/task/continuous-control","name":"Continuous Control","papers":6},{"task":"/task/continuous-control","name":"continuous-control","papers":6},{"task":"/task/openai-gym","name":"OpenAI Gym","papers":5},{"task":"/task/benchmarking","name":"Benchmarking","papers":4},{"task":"/task/management","name":"Management","papers":4},{"task":"/task/representation-learning","name":"Representation Learning","papers":4},{"task":null,"name":"GPU","papers":3},{"task":"/task/mujoco","name":"MuJoCo","papers":3},{"task":"/task/multi-agent-reinforcement-learning","name":"Multi-agent Reinforcement Learning","papers":3},{"task":"/task/computational-efficiency","name":"Computational Efficiency","papers":2},{"task":"/task/drug-discovery","name":"Drug Discovery","papers":2},{"task":"/task/language-modelling","name":"Language Modelling","papers":2},{"task":"/task/myocardial-infarction-detection","name":"Myocardial infarction detection","papers":2}],"tasks_shown":20,"n_tasks":86,"usage_by_year":[{"year":"2016","papers":1},{"year":"2017","papers":4},{"year":"2018","papers":8},{"year":"2019","papers":14},{"year":"2020","papers":10},{"year":"2021","papers":8},{"year":"2022","papers":10},{"year":"2023","papers":14},{"year":"2024","papers":9},{"year":"2025","papers":4}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/a2c"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}