Methods › Reinforcement Learning › Policy Gradient Methods › A2C
A2C
Introduced by Volodymyr Mnih et al. in Asynchronous Methods for Deep Reinforcement Learning
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
A2C, or Advantage Actor Critic, is a synchronous version of the A3C policy gradient method. As an alternative to the asynchronous implementation of A3C, A2C is a synchronous, deterministic implementation that waits for each actor to finish its segment of experience before updating, averaging over all of the actors. This more effectively uses GPUs due to larger batch sizes.
Image Credit: OpenAI Baselines
Papers archive 2025-07-28
30 shown of 82, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control 13 May 2025 · 0 repositories · arXiv:2505.09029
-
Ensemble RL through Classifier Models: Enhancing Risk-Return Trade-offs in Trading Strategies 23 Feb 2025 · 0 repositories · arXiv:2502.17518
-
Continuous Learning Conversational AI: A Personalized Agent Framework via A2C Reinforcement Learning 18 Feb 2025 · 0 repositories · arXiv:2502.12876
-
EvoRL: A GPU-accelerated Framework for Evolutionary Reinforcement Learning 25 Jan 2025 · 2 repositories · arXiv:2501.15129
-
Innate-Values-driven Reinforcement Learning based Cognitive Modeling 14 Nov 2024 · 0 repositories · arXiv:2411.09160
-
Deep Reinforcement Learning Strategies in Finance: Insights into Asset Holding, Trading Behavior, and Purchase Diversity 29 Jun 2024 · 0 repositories · arXiv:2407.09557
-
Multistep Criticality Search and Power Shaping in Microreactors with Reinforcement Learning 22 Jun 2024 · 0 repositories · arXiv:2406.15931
-
Biological Neurons Compete with Deep Reinforcement Learning in Sample Efficiency in a Simulated Gameworld 27 May 2024 · 0 repositories · arXiv:2405.16946
-
Symmetric Reinforcement Learning Loss for Robust Learning on Diverse Tasks and Model Scales 27 May 2024 · 1 repository · arXiv:2405.17618
-
Extracting Heuristics from Large Language Models for Reward Shaping in Reinforcement Learning 24 May 2024 · 0 repositories · arXiv:2405.15194
-
Portfolio Management using Deep Reinforcement Learning 1 May 2024 · 0 repositories · arXiv:2405.01604
-
Breaching the Bottleneck: Evolutionary Transition from Reward-Driven Learning to Reward-Agnostic Domain-Adapted Learning in Neuromodulated Neural Nets 19 Apr 2024 · 1 repository · arXiv:2404.12631
-
A2C: A Modular Multi-stage Collaborative Decision Framework for Human-AI Teams 25 Jan 2024 · 0 repositories · arXiv:2401.14432
-
Epidemic Decision-making System Based Federated Reinforcement Learning 3 Nov 2023 · 0 repositories · arXiv:2311.01749
-
Diagnosis-oriented Medical Image Compression with Efficient Transfer Learning 20 Oct 2023 · 0 repositories · arXiv:2310.13250
-
Deep Reinforcement Learning-based Intelligent Traffic Signal Controls with Optimized CO2 emissions 19 Oct 2023 · 1 repository · arXiv:2310.13129
-
Solving the Quadratic Assignment Problem using Deep Reinforcement Learning 2 Oct 2023 · 0 repositories · arXiv:2310.01604
-
Raijū: Reinforcement Learning-Guided Post-Exploitation for Automating Security Assessment of Network Systems 27 Sep 2023 · 0 repositories · arXiv:2309.15518
-
SAF-Net: Self-Attention Fusion Network for Myocardial Infarction Detection using Multi-View Echocardiography 27 Sep 2023 · 0 repositories · arXiv:2309.15520
-
Career Path Recommendations for Long-term Income Maximization: A Reinforcement Learning Approach 11 Sep 2023 · 0 repositories · arXiv:2309.05391
-
Semantic Consistency for Assuring Reliability of Large Language Models 17 Aug 2023 · 0 repositories · arXiv:2308.09138
-
Deep Reinforcement Learning for ESG financial portfolio management 19 Jun 2023 · 0 repositories · arXiv:2307.09631
-
Multi-Agent Reinforcement Learning for Network Routing in Integrated Access Backhaul Networks 12 May 2023 · 0 repositories · arXiv:2305.16170
-
Deep reinforcement learning applied to an assembly sequence planning problem with user preferences 13 Apr 2023 · 0 repositories · arXiv:2304.06567
-
Read and Reap the Rewards: Learning to Play Atari with the Help of Instruction Manuals 9 Feb 2023 · 0 repositories · arXiv:2302.04449
-
A Scale-Independent Multi-Objective Reinforcement Learning with Convergence Analysis 8 Feb 2023 · 0 repositories · arXiv:2302.04179
-
Learning, Fast and Slow: A Goal-Directed Memory-Based Approach for Dynamic Environments 31 Jan 2023 · 1 repository · arXiv:2301.13758
-
Towards automating Codenames spymasters with deep reinforcement learning 28 Dec 2022 · 0 repositories · arXiv:2212.14104
-
Reinforcement Learning for Molecular Dynamics Optimization: A Stochastic Pontryagin Maximum Principle Approach 6 Dec 2022 · 1 repository · arXiv:2212.03320
-
Towards More Efficient Shared Autonomous Mobility: A Learning-Based Fleet Repositioning Approach 16 Oct 2022 · 0 repositories · arXiv:2210.08659
Tasks archive 2025-07-28
20 shown of 86 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections