Methods › Reinforcement Learning › Off-Policy TD Control › Q-Learning › Papers, page 18
Q-Learning
Papers archive 2025-07-28
archive papers tagged: 1,734 · with a code link: 464 · where Syntology ran a sample: 126 (105 with a run with no instrument failure, 21 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (126 of 1,734 tagged: 105 with a run with no instrument failure, 21 where every run was a failure of Syntology's instrument)
Page 18 of 18: papers 1,701 to 1,734 of 1,734, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Deep Exploration via Bootstrapped DQN 15 Feb 2016 · 6 repositories · arXiv:1602.04621Syntology 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Using Deep Q-Learning to Control Optimization Hyperparameters 12 Feb 2016 · 0 repositories · arXiv:1602.04062
-
Angrier Birds: Bayesian reinforcement learning 6 Jan 2016 · 1 repository · arXiv:1601.01297
-
How to Discount Deep Reinforcement Learning: Towards New Dynamic Strategies 7 Dec 2015 · 0 repositories · arXiv:1512.02011
-
Deep Attention Recurrent Q-Network 5 Dec 2015 · 3 repositories · arXiv:1512.01693
-
State of the Art Control of Atari Games Using Shallow Reinforcement Learning 4 Dec 2015 · 1 repository · arXiv:1512.01563
-
Deep Reinforcement Learning with Attention for Slate Markov Decision Processes with High-Dimensional States and Actions 3 Dec 2015 · 0 repositories · arXiv:1512.01124
-
Multiagent Cooperation and Competition with Deep Reinforcement Learning 27 Nov 2015 · 4 repositories · arXiv:1511.08779
-
Policy Distillation 19 Nov 2015 · 1 repository · arXiv:1511.06295
-
Prioritized Experience Replay 18 Nov 2015 · 77 repositories · arXiv:1511.05952Syntology 80 ran (of which 62 constructed an object rather than computing a result; 72 with no instrument failure: 4 honoured, 0 violated, 68 with no contract checked; 8 where Syntology's instrument failed) · 31 unverified (of 111 harvested samples) · 43 pointer-only (licence)
-
Deep Reinforcement Learning with a Natural Language Action Space 14 Nov 2015 · 3 repositories · arXiv:1511.04636
-
A disembodied developmental robotic agent called Samu Bátfai 9 Nov 2015 · 15 repositories · arXiv:1511.02889
-
Generating Text with Deep Reinforcement Learning 30 Oct 2015 · 0 repositories · arXiv:1510.09202
-
Deep Reinforcement Learning with Double Q-learning 22 Sep 2015 · 97 repositories · arXiv:1509.06461Syntology 56 ran (of which 38 constructed an object rather than computing a result; 55 with no instrument failure: 0 honoured, 0 violated, 55 with no contract checked; 1 where Syntology's instrument failed) · 50 unverified (of 106 harvested samples) · 57 pointer-only (licence)
-
Optimization of anemia treatment in hemodialysis patients via reinforcement learning 14 Sep 2015 · 0 repositories · arXiv:1509.03977
-
Continuous control with deep reinforcement learning 9 Sep 2015 · 161 repositories · arXiv:1509.02971Syntology 163 ran (of which 126 constructed an object rather than computing a result; 152 with no instrument failure: 3 honoured, 0 violated, 149 with no contract checked; 11 where Syntology's instrument failed) · 143 unverified (of 306 harvested samples) · 163 pointer-only (licence)
-
Artificial Prediction Markets for Online Prediction of Continuous Variables-A Preliminary Report 11 Aug 2015 · 0 repositories · arXiv:1508.02681
-
Massively Parallel Methods for Deep Reinforcement Learning 15 Jul 2015 · 3 repositories · arXiv:1507.04296
-
Online Transfer Learning in Reinforcement Learning Domains 2 Jul 2015 · 0 repositories · arXiv:1507.00436
-
Autonomous CRM Control via CLV Approximation with Deep Reinforcement Learning in Discrete and Continuous Action Space 8 Apr 2015 · 0 repositories · arXiv:1504.01840
-
Deep Learning for Real-Time Atari Game Play Using Offline Monte-Carlo Tree Search Planning 1 Dec 2014 · 0 repositories
-
Empirical Q-Value Iteration 30 Nov 2014 · 0 repositories · arXiv:1412.0180
-
Q-learning for Optimal Control of Continuous-time Systems 11 Oct 2014 · 0 repositories · arXiv:1410.2954
-
Reinforcement Learning Based Algorithm for the Maximization of EV Charging Station Revenue 4 Jul 2014 · 0 repositories · arXiv:1407.1291
-
Personalized Medical Treatments Using Novel Reinforcement Learning Algorithms 16 Jun 2014 · 0 repositories · arXiv:1406.3922
-
Two Timescale Convergent Q-learning for Sleep--Scheduling in Wireless Sensor Networks 27 Dec 2013 · 0 repositories · arXiv:1312.7292
-
Playing Atari with Deep Reinforcement Learning 19 Dec 2013 · 112 repositories · arXiv:1312.5602Syntology 64 ran (of which 24 constructed an object rather than computing a result; 46 with no instrument failure: 5 honoured, 0 violated, 41 with no contract checked; 18 where Syntology's instrument failed) · 53 unverified (of 117 harvested samples) · 56 pointer-only (licence)
-
Q-learning optimization in a multi-agents system for image segmentation 23 Nov 2013 · 0 repositories · arXiv:1311.6054
-
Risk-sensitive Reinforcement Learning 8 Nov 2013 · 0 repositories · arXiv:1311.2097
-
Approximate Kalman Filter Q-Learning for Continuous State-Space MDPs 26 Sep 2013 · 0 repositories · arXiv:1309.6868
-
Projective simulation for classical learning agents: a comprehensive investigation 7 May 2013 · 0 repositories · arXiv:1305.1578
-
Speedy Q-Learning 1 Dec 2011 · 0 repositories
-
Double Q-learning 1 Dec 2010 · 0 repositories
-
Convergent Temporal-Difference Learning with Arbitrary Smooth Function Approximation 1 Dec 2009 · 0 repositories