Methods › Reinforcement Learning › Off-Policy TD Control › Q-Learning › Papers where code ran, page 2
Q-Learning
Papers archive 2025-07-28
archive papers tagged: 1,734 · with a code link: 464 · where Syntology ran a sample: 126 (105 with a run with no instrument failure, 21 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (126 of 1,734 tagged: 105 with a run with no instrument failure, 21 where every run was a failure of Syntology's instrument)
Syntology We ran code from the paper's repository; we did not isolate this method inside it.
Page 2 of 2: papers 101 to 126 of the 126 tagged papers where Syntology ran at least one harvested sample (105 with a run with no instrument failure, 21 where every run was a failure of Syntology's instrument), newest first by the archive's date (ties by arXiv id). This is a filter on Syntology's record ordered by date only, not a ranking; a run is not a correctness claim. A paper missing from this list is not a recorded non-run: it may have no arXiv id, no harvested code, or only samples that have not run yet.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Off-Policy Deep Reinforcement Learning without Exploration 7 Dec 2018 · 10 repositories · arXiv:1812.02900Syntology community repositories only · 14 ran (of which 12 constructed an object rather than computing a result; 14 with no instrument failure: 1 honoured, 0 violated, 13 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 14 harvested samples) · 9 pointer-only (licence)
-
Deep Multi-Agent Reinforcement Learning with Relevance Graphs 30 Nov 2018 · 1 repository · arXiv:1811.12557Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 6 unverified (of 7 harvested samples)
-
Greedy Actor-Critic: A New Conditional Cross-Entropy Method for Policy Improvement 22 Oct 2018 · 1 repository · arXiv:1810.09103Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 10 harvested samples)
-
Successor Uncertainties: Exploration and Uncertainty in Temporal Difference Learning 15 Oct 2018 · 2 repositories · arXiv:1810.06530Syntology 7 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 1 violated, 0 with no contract checked; 5 where Syntology's instrument failed) · 0 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
Parametrized Deep Q-Networks Learning: Reinforcement Learning with Discrete-Continuous Hybrid Action Space 10 Oct 2018 · 5 repositories · arXiv:1810.06394Syntology 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Towards Better Interpretability in Deep Q-Networks 15 Sep 2018 · 1 repository · arXiv:1809.05630Syntology 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Implicit Quantile Networks for Distributional Reinforcement Learning 14 Jun 2018 · 19 repositories · arXiv:1806.06923Syntology 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Learning Synergies between Pushing and Grasping with Self-supervised Deep Reinforcement Learning 27 Mar 2018 · 4 repositories · arXiv:1803.09956Syntology official (archive's flag): 3 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 11 harvested samples)
-
Mean Field Multi-Agent Reinforcement Learning 15 Feb 2018 · 3 repositories · arXiv:1802.05438Syntology official (archive's flag): 1 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 2 pointer-only (licence)
-
Improving Exploration in Evolution Strategies for Deep Reinforcement Learning via a Population of Novelty-Seeking Agents 18 Dec 2017 · 2 repositories · arXiv:1712.06560Syntology official: no sample here; runs from other or unrecorded repositories · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Deep Neuroevolution: Genetic Algorithms Are a Competitive Alternative for Training Deep Neural Networks for Reinforcement Learning 18 Dec 2017 · 12 repositories · arXiv:1712.06567Syntology 4 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 4 pointer-only (licence)
-
TreeQN and ATreeC: Differentiable Tree-Structured Models for Deep Reinforcement Learning 31 Oct 2017 · 1 repository · arXiv:1710.11417Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 12 harvested samples) · 1 pointer-only (licence)
-
Distributional Reinforcement Learning with Quantile Regression 27 Oct 2017 · 17 repositories · arXiv:1710.10044Syntology 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Rainbow: Combining Improvements in Deep Reinforcement Learning 6 Oct 2017 · 34 repositories · arXiv:1710.02298Syntology 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 1 honoured, 1 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples) · 1 pointer-only (licence)
-
Noisy Networks for Exploration 30 Jun 2017 · 15 repositories · arXiv:1706.10295Syntology 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments 7 Jun 2017 · 86 repositories · arXiv:1706.02275Syntology official: no sample here; runs from other or unrecorded repositories · 75 ran (of which 54 constructed an object rather than computing a result; 68 with no instrument failure: 2 honoured, 0 violated, 66 with no contract checked; 7 where Syntology's instrument failed) · 68 unverified (of 143 harvested samples) · 99 pointer-only (licence)
-
Parameter Space Noise for Exploration 6 Jun 2017 · 10 repositories · arXiv:1706.01905Syntology 4 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Evolution Strategies as a Scalable Alternative to Reinforcement Learning 10 Mar 2017 · 23 repositories · arXiv:1703.03864Syntology official (archive's flag): 4 ran · 16 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 1 violated, 10 with no contract checked; 5 where Syntology's instrument failed) · 13 unverified (of 29 harvested samples) · 2 pointer-only (licence)
-
Count-Based Exploration with Neural Density Models 3 Mar 2017 · 1 repository · arXiv:1703.01310Syntology 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
ViZDoom: A Doom-based AI Research Platform for Visual Reinforcement Learning 6 May 2016 · 10 repositories · arXiv:1605.02097Syntology community repositories only · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Continuous Deep Q-Learning with Model-based Acceleration 2 Mar 2016 · 8 repositories · arXiv:1603.00748Syntology 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Deep Exploration via Bootstrapped DQN 15 Feb 2016 · 6 repositories · arXiv:1602.04621Syntology 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Prioritized Experience Replay 18 Nov 2015 · 77 repositories · arXiv:1511.05952Syntology 80 ran (of which 62 constructed an object rather than computing a result; 72 with no instrument failure: 4 honoured, 0 violated, 68 with no contract checked; 8 where Syntology's instrument failed) · 31 unverified (of 111 harvested samples) · 43 pointer-only (licence)
-
Deep Reinforcement Learning with Double Q-learning 22 Sep 2015 · 97 repositories · arXiv:1509.06461Syntology 56 ran (of which 38 constructed an object rather than computing a result; 55 with no instrument failure: 0 honoured, 0 violated, 55 with no contract checked; 1 where Syntology's instrument failed) · 50 unverified (of 106 harvested samples) · 57 pointer-only (licence)
-
Continuous control with deep reinforcement learning 9 Sep 2015 · 161 repositories · arXiv:1509.02971Syntology 163 ran (of which 126 constructed an object rather than computing a result; 152 with no instrument failure: 3 honoured, 0 violated, 149 with no contract checked; 11 where Syntology's instrument failed) · 143 unverified (of 306 harvested samples) · 163 pointer-only (licence)
-
Playing Atari with Deep Reinforcement Learning 19 Dec 2013 · 112 repositories · arXiv:1312.5602Syntology 64 ran (of which 24 constructed an object rather than computing a result; 46 with no instrument failure: 5 honoured, 0 violated, 41 with no contract checked; 18 where Syntology's instrument failed) · 53 unverified (of 117 harvested samples) · 56 pointer-only (licence)