Methods › Reinforcement Learning › Board Game Models › AlphaZero › Papers where code ran, page 1
AlphaZero
Papers archive 2025-07-28
archive papers tagged: 114 · with a code link: 60 · where Syntology ran a sample: 17 (16 with a run with no instrument failure, 1 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (17 of 114 tagged: 16 with a run with no instrument failure, 1 where every run was a failure of Syntology's instrument)
Syntology We ran code from the paper's repository; we did not isolate this method inside it.
Page 1 of 1: papers 1 to 17 of the 17 tagged papers where Syntology ran at least one harvested sample (16 with a run with no instrument failure, 1 where every run was a failure of Syntology's instrument), newest first by the archive's date (ties by arXiv id). This is a filter on Syntology's record ordered by date only, not a ranking; a run is not a correctness claim. A paper missing from this list is not a recorded non-run: it may have no arXiv id, no harvested code, or only samples that have not run yet.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
On the Emergence of Thinking in LLMs I: Searching for the Right Intuition 10 Feb 2025 · 4 repositories · arXiv:2502.06773Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 13 harvested samples) · 9 pointer-only (licence)
-
AlphaZero Neural Scaling and Zipf's Law: a Tale of Board Games and Power Laws 16 Dec 2024 · 1 repository · arXiv:2412.11979Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 8 harvested samples)
-
Bayes Adaptive Monte Carlo Tree Search for Offline Model-based Reinforcement Learning 15 Oct 2024 · 1 repository · arXiv:2410.11234Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Maia-2: A Unified Model for Human-AI Alignment in Chess 30 Sep 2024 · 2 repositories · arXiv:2409.20553Syntology official (archive's flag): 4 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 10 unverified (of 18 harvested samples) · 11 pointer-only (licence)
-
Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning 1 May 2024 · 2 repositories · arXiv:2405.00451Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 1 pointer-only (licence)
-
SPO: Sequential Monte Carlo Policy Optimisation 12 Feb 2024 · 1 repository · arXiv:2402.07963Syntology 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples)
-
Amortized Planning with Large-Scale Transformers: A Case Study on Chess 7 Feb 2024 · 1 repository · arXiv:2402.04494Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples)
-
Game Solving with Online Fine-Tuning 13 Nov 2023 · 1 repository · arXiv:2311.07178Syntology official (archive's flag): 5 ran · 5 ran (of which 5 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; every one of the 5 samples that ran constructed an object rather than computing a result (of 5 harvested samples) · 5 pointer-only (licence)
-
MiniZero: Comparative Analysis of AlphaZero and MuZero on Go, Othello, and Atari Games 17 Oct 2023 · 1 repository · arXiv:2310.11305Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples)
-
Information based explanation methods for deep learning agents -- with applications on large open-source chess models 18 Sep 2023 · 1 repository · arXiv:2309.09702Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Evaluation Beyond Task Performance: Analyzing Concepts in AlphaZero in Hex 26 Nov 2022 · 1 repository · arXiv:2211.14673Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
Scaling Laws for a Multi-Agent Reinforcement Learning Model 29 Sep 2022 · 1 repository · arXiv:2210.00849Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Neural Networks for Chess 3 Sep 2022 · 2 repositories · arXiv:2209.01506Syntology community repositories only · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
Scaling Scaling Laws with Board Games 7 Apr 2021 · 2 repositories · arXiv:2104.03113Syntology 13 ran (of which 0 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 0 violated, 13 with no contract checked; 0 where Syntology's instrument failed) · 7 unverified (of 20 harvested samples)
-
Combining Deep Reinforcement Learning and Search for Imperfect-Information Games 27 Jul 2020 · 1 repository · arXiv:2007.13544Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model 19 Nov 2019 · 18 repositories · arXiv:1911.08265Syntology 43 ran (of which 36 constructed an object rather than computing a result; 43 with no instrument failure: 4 honoured, 0 violated, 39 with no contract checked; 0 where Syntology's instrument failed) · 21 unverified (of 64 harvested samples) · 62 pointer-only (licence)
-
Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm 5 Dec 2017 · 62 repositories · arXiv:1712.01815Syntology 13 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 3 honoured, 1 violated, 0 with no contract checked; 9 where Syntology's instrument failed) · 4 unverified (of 17 harvested samples) · 9 pointer-only (licence)