Methods › Reinforcement Learning › Off-Policy TD Control › Double Q-learning › Papers, page 2
Double Q-learning
Papers archive 2025-07-28
archive papers tagged: 112 · with a code link: 45 · where Syntology ran a sample: 15 (13 with a run with no instrument failure, 2 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (15 of 112 tagged: 13 with a run with no instrument failure, 2 where every run was a failure of Syntology's instrument)
Page 2 of 2: papers 101 to 112 of 112, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Distributed Prioritized Experience Replay 2 Mar 2018 · 15 repositories · arXiv:1803.00933Syntology 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 6 unverified (of 15 harvested samples) · 3 pointer-only (licence)
-
Addressing Function Approximation Error in Actor-Critic Methods 26 Feb 2018 · 67 repositories · arXiv:1802.09477Syntology community repositories only · 26 ran (of which 0 constructed an object rather than computing a result; 25 with no instrument failure: 1 honoured, 1 violated, 23 with no contract checked; 1 where Syntology's instrument failed) · 10 unverified (of 36 harvested samples) · 21 pointer-only (licence)
-
Weighted Double Deep Multiagent Reinforcement Learning in Stochastic Cooperative Environments 23 Feb 2018 · 0 repositories · arXiv:1802.08534
-
Efficient Exploration through Bayesian Deep Q-Networks 13 Feb 2018 · 1 repository · arXiv:1802.04412
-
Faster Deep Q-learning using Neural Episodic Control 6 Jan 2018 · 0 repositories · arXiv:1801.01968
-
Rainbow: Combining Improvements in Deep Reinforcement Learning 6 Oct 2017 · 34 repositories · arXiv:1710.02298Syntology 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 1 honoured, 1 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples) · 1 pointer-only (licence)
-
Noisy Networks for Exploration 30 Jun 2017 · 15 repositories · arXiv:1706.10295Syntology 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Sample Efficient Actor-Critic with Experience Replay 3 Nov 2016 · 7 repositories · arXiv:1611.01224Syntology 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Dynamic Frame skip Deep Q Network 17 May 2016 · 0 repositories · arXiv:1605.05365
-
Dueling Network Architectures for Deep Reinforcement Learning 20 Nov 2015 · 73 repositories · arXiv:1511.06581Syntology 7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 3 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified (of 11 harvested samples) · 7 pointer-only (licence)
-
Deep Reinforcement Learning with Double Q-learning 22 Sep 2015 · 97 repositories · arXiv:1509.06461Syntology 56 ran (of which 38 constructed an object rather than computing a result; 55 with no instrument failure: 0 honoured, 0 violated, 55 with no contract checked; 1 where Syntology's instrument failed) · 50 unverified (of 106 harvested samples) · 57 pointer-only (licence)
-
Double Q-learning 1 Dec 2010 · 0 repositories