Methods › Reinforcement Learning › Board Game Models › AlphaZero
AlphaZero
Introduced by David Silver et al. in Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
AlphaZero is a reinforcement learning agent for playing board games such as Go, chess, and shogi.
Papers archive 2025-07-28
30 shown of 114, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
AlphaZero-Edu: Making AlphaZero Accessible to Everyone 20 Apr 2025 · 1 repository · arXiv:2504.14636
-
AssistanceZero: Scalably Solving Assistance Games 9 Apr 2025 · 1 repository · arXiv:2504.07091
-
Reinforcement Learning and Life Cycle Assessment for a Circular Economy -- Towards Progressive Computer Science 13 Mar 2025 · 0 repositories · arXiv:2503.10822
-
Alignment, Agency and Autonomy in Frontier AI: A Systems Engineering Perspective 20 Feb 2025 · 0 repositories · arXiv:2503.05748
-
Playing Hex and Counter Wargames using Reinforcement Learning and Recurrent Neural Networks 19 Feb 2025 · 1 repository · arXiv:2502.13918
-
On the Emergence of Thinking in LLMs I: Searching for the Right Intuition 10 Feb 2025 · 4 repositories · arXiv:2502.06773Syntology ran 0 of 13 samples · 13 unverified
-
Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning 23 Dec 2024 · 0 repositories · arXiv:2412.17397
-
AlphaZero Neural Scaling and Zipf's Law: a Tale of Board Games and Power Laws 16 Dec 2024 · 1 repository · arXiv:2412.11979Syntology ran 0 of 8 samples · 8 unverified
-
Mastering NIM and Impartial Games with Weak Neural Networks: An AlphaZero-inspired Multi-Frame Approach 10 Nov 2024 · 0 repositories · arXiv:2411.06403
-
Enhancing Chess Reinforcement Learning with Graph Representation 31 Oct 2024 · 1 repository · arXiv:2410.23753
-
Bayes Adaptive Monte Carlo Tree Search for Offline Model-based Reinforcement Learning 15 Oct 2024 · 1 repository · arXiv:2410.11234Syntology ran 2 of 2 samples · 0 unverified
-
ResTNet: Defense against Adversarial Policies via Transformer in Computer Go 7 Oct 2024 · 0 repositories · arXiv:2410.05347
-
Maia-2: A Unified Model for Human-AI Alignment in Chess 30 Sep 2024 · 2 repositories · arXiv:2409.20553Syntology ran 8 of 18 samples · 10 unverified · 11 pointer-only (licence)
-
Mastering Chess with a Transformer Model 18 Sep 2024 · 1 repository · arXiv:2409.12272
-
AlphaViT: A Flexible Game-Playing AI for Multiple Games and Variable Board Sizes 25 Aug 2024 · 1 repository · arXiv:2408.13871
-
ShortCircuit: AlphaZero-Driven Circuit Design 19 Aug 2024 · 0 repositories · arXiv:2408.09858
-
Structure and Reduction of MCTS for Explainable-AI 10 Aug 2024 · 0 repositories · arXiv:2408.05488
-
Provably Efficient Long-Horizon Exploration in Monte Carlo Tree Search through State Occupancy Regularization 7 Jul 2024 · 0 repositories · arXiv:2407.05511
-
AlphaZeroES: Direct score maximization outperforms planning loss minimization 12 Jun 2024 · 0 repositories · arXiv:2406.08687
-
Learning to Play 7 Wonders Duel Without Human Supervision 2 Jun 2024 · 0 repositories · arXiv:2406.00741
-
Model Predictive Control and Reinforcement Learning: A Unified Framework Based on Dynamic Programming 2 Jun 2024 · 0 repositories · arXiv:2406.00592
-
Super-Exponential Regret for UCT, AlphaGo and Variants 7 May 2024 · 0 repositories · arXiv:2405.04407
-
Model-based reinforcement learning for protein backbone design 3 May 2024 · 0 repositories · arXiv:2405.01983
-
Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning 1 May 2024 · 2 repositories · arXiv:2405.00451Syntology ran 1 of 2 samples · 1 unverified
-
Policy Mirror Descent with Lookahead 21 Mar 2024 · 1 repository · arXiv:2403.14156
-
SPO: Sequential Monte Carlo Policy Optimisation 12 Feb 2024 · 1 repository · arXiv:2402.07963Syntology ran 0 of 5 samples · 5 unverified
-
Amortized Planning with Large-Scale Transformers: A Case Study on Chess 7 Feb 2024 · 1 repository · arXiv:2402.04494Syntology ran 5 of 7 samples · 2 unverified
-
Mastering Zero-Shot Interactions in Cooperative and Competitive Simultaneous Games 5 Feb 2024 · 0 repositories · arXiv:2402.03136
-
Discovering Mathematical Formulas from Data via GPT-guided Monte Carlo Tree Search 24 Jan 2024 · 0 repositories · arXiv:2401.14424
-
Decision Making in Non-Stationary Environments with Policy-Augmented Search 6 Jan 2024 · 1 repository · arXiv:2401.03197
Tasks archive 2025-07-28
20 shown of 65 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections