Methods › Reinforcement Learning › Off-Policy TD Control › Q-Learning › Papers, page 17
Q-Learning
Papers archive 2025-07-28
archive papers tagged: 1,734 · with a code link: 464 · where Syntology ran a sample: 126 (105 with a run with no instrument failure, 21 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (126 of 1,734 tagged: 105 with a run with no instrument failure, 21 where every run was a failure of Syntology's instrument)
Page 17 of 18: papers 1,601 to 1,700 of 1,734, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Monte Carlo Q-learning for General Game Playing 16 Feb 2018 · 2 repositories · arXiv:1802.05944
-
Mean Field Multi-Agent Reinforcement Learning 15 Feb 2018 · 3 repositories · arXiv:1802.05438Syntology official (archive's flag): 1 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 2 pointer-only (licence)
-
Efficient Exploration through Bayesian Deep Q-Networks 13 Feb 2018 · 1 repository · arXiv:1802.04412
-
Q-learning with Nearest Neighbors 12 Feb 2018 · 0 repositories · arXiv:1802.03900
-
Balancing Two-Player Stochastic Games with Soft Q-Learning 9 Feb 2018 · 0 repositories · arXiv:1802.03216
-
Deep Reinforcement Learning using Capsules in Advanced Game Environments 29 Jan 2018 · 0 repositories · arXiv:1801.09597
-
The QLBS Q-Learner Goes NuQLear: Fitted Q Iteration, Inverse RL, and Option Portfolios 17 Jan 2018 · 0 repositories · arXiv:1801.06077
-
Deep Reinforcement Fuzzing 14 Jan 2018 · 0 repositories · arXiv:1801.04589
-
DeepTraffic: Crowdsourced Hyperparameter Tuning of Deep Reinforcement Learning Systems for Multi-Agent Dense Traffic Navigation 9 Jan 2018 · 6 repositories · arXiv:1801.02805
-
Faster Deep Q-learning using Neural Episodic Control 6 Jan 2018 · 0 repositories · arXiv:1801.01968
-
Autonomous Vehicle Fleet Coordination With Deep Reinforcement Learning 1 Jan 2018 · 0 repositories
-
Avoiding Catastrophic States with Intrinsic Fear 1 Jan 2018 · 0 repositories
-
Faster Reinforcement Learning with Expert State Sequences 1 Jan 2018 · 0 repositories
-
PARAMETRIZED DEEP Q-NETWORKS LEARNING: PLAYING ONLINE BATTLE ARENA WITH DISCRETE-CONTINUOUS HYBRID ACTION SPACE 1 Jan 2018 · 1 repository
-
Representing Entropy : A short proof of the equivalence between soft Q-learning and policy gradients 1 Jan 2018 · 0 repositories
-
TD Learning with Constrained Gradients 1 Jan 2018 · 0 repositories
-
A short variational proof of equivalence between policy gradients and soft Q learning 22 Dec 2017 · 0 repositories · arXiv:1712.08650
-
A Deep Policy Inference Q-Network for Multi-Agent Systems 21 Dec 2017 · 0 repositories · arXiv:1712.07893
-
Deep Neuroevolution: Genetic Algorithms Are a Competitive Alternative for Training Deep Neural Networks for Reinforcement Learning 18 Dec 2017 · 12 repositories · arXiv:1712.06567Syntology 4 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 4 pointer-only (licence)
-
Improving Exploration in Evolution Strategies for Deep Reinforcement Learning via a Population of Novelty-Seeking Agents 18 Dec 2017 · 2 repositories · arXiv:1712.06560Syntology official: no sample here; runs from other or unrecorded repositories · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Towards a Deep Reinforcement Learning Approach for Tower Line Wars 17 Dec 2017 · 0 repositories · arXiv:1712.06180
-
QLBS: Q-Learner in the Black-Scholes(-Merton) Worlds 13 Dec 2017 · 1 repository · arXiv:1712.04609
-
Assumed Density Filtering Q-learning 9 Dec 2017 · 1 repository · arXiv:1712.03333
-
Deep Primal-Dual Reinforcement Learning: Accelerating Actor-Critic using Bellman Duality 7 Dec 2017 · 0 repositories · arXiv:1712.02467
-
Q-LDA: Uncovering Latent Patterns in Text-based Sequential Decision Processes 1 Dec 2017 · 0 repositories
-
Zap Q-Learning 1 Dec 2017 · 0 repositories
-
Uncertainty Estimates for Efficient Neural Network-based Dialogue Policy Optimisation 30 Nov 2017 · 0 repositories · arXiv:1711.11486
-
A Benchmarking Environment for Reinforcement Learning Based Task Oriented Dialogue Management 29 Nov 2017 · 0 repositories · arXiv:1711.11023
-
Implementing the Deep Q-Network 20 Nov 2017 · 1 repository · arXiv:1711.07478
-
Neural Network Based Reinforcement Learning for Audio-Visual Gaze Control in Human-Robot Interaction 18 Nov 2017 · 0 repositories · arXiv:1711.06834
-
BBQ-Networks: Efficient Exploration in Deep Reinforcement Learning for Task-Oriented Dialogue Systems 15 Nov 2017 · 0 repositories · arXiv:1711.05715
-
A unified decision making framework for supply and demand management in microgrid networks 14 Nov 2017 · 0 repositories · arXiv:1711.05078
-
Double Q(σ) and Q(σ, λ): Unifying Reinforcement Learning Control Algorithms 5 Nov 2017 · 0 repositories · arXiv:1711.01569
-
Adaptive coordination of working-memory and reinforcement learning in non-human primates performing a trial-and-error problem solving task 2 Nov 2017 · 1 repository · arXiv:1711.00698
-
TreeQN and ATreeC: Differentiable Tree-Structured Models for Deep Reinforcement Learning 31 Oct 2017 · 1 repository · arXiv:1710.11417Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 12 harvested samples) · 1 pointer-only (licence)
-
Distributional Reinforcement Learning with Quantile Regression 27 Oct 2017 · 17 repositories · arXiv:1710.10044Syntology 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
The Effects of Memory Replay in Reinforcement Learning 18 Oct 2017 · 1 repository · arXiv:1710.06574
-
Deep Reinforcement Learning: Framework, Applications, and Embedded Implementations 10 Oct 2017 · 0 repositories · arXiv:1710.03792
-
Rainbow: Combining Improvements in Deep Reinforcement Learning 6 Oct 2017 · 34 repositories · arXiv:1710.02298Syntology 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 1 honoured, 1 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples) · 1 pointer-only (licence)
-
Supervised Q-walk for Learning Vector Representation of Nodes in Networks 3 Oct 2017 · 0 repositories · arXiv:1710.00978
-
A Simple Reinforcement Learning Mechanism for Resource Allocation in LTE-A Networks with Markov Decision Process and Q-Learning 27 Sep 2017 · 0 repositories · arXiv:1709.09312
-
An Optimal Online Method of Selecting Source Policies for Reinforcement Learning 24 Sep 2017 · 0 repositories · arXiv:1709.08201
-
Improving Search through A3C Reinforcement Learning based Conversational Agent 17 Sep 2017 · 0 repositories · arXiv:1709.05638
-
Deep Reinforcement Learning with Surrogate Agent-Environment Interface 12 Sep 2017 · 0 repositories · arXiv:1709.03942
-
Pre-training Neural Networks with Human Demonstrations for Deep Reinforcement Learning 12 Sep 2017 · 0 repositories · arXiv:1709.04083
-
Formulation of Deep Reinforcement Learning Architecture Toward Autonomous Driving for On-Ramp Merge 7 Sep 2017 · 0 repositories · arXiv:1709.02066
-
BIBI System Description: Building with CNNs and Breaking with Deep Reinforcement Learning 1 Sep 2017 · 0 repositories
-
Multi-Agent Q-Learning for Minimizing Demand-Supply Power Deficit in Microgrids 25 Aug 2017 · 0 repositories · arXiv:1708.07732
-
LADDER: A Human-Level Bidding Agent for Large-Scale Real-Time Online Auctions 18 Aug 2017 · 0 repositories · arXiv:1708.05565
-
Practical Block-wise Neural Network Architecture Generation 18 Aug 2017 · 1 repository · arXiv:1708.05552
-
Investigating Reinforcement Learning Agents for Continuous State Space Environments 8 Aug 2017 · 0 repositories · arXiv:1708.02378
-
3DCNN-DQN-RNN: A Deep Reinforcement Learning Framework for Semantic Parsing of Large-scale 3D Point Clouds 21 Jul 2017 · 0 repositories · arXiv:1707.06783
-
Empirical evaluation of a Q-Learning Algorithm for Model-free Autonomous Soaring 18 Jul 2017 · 0 repositories · arXiv:1707.05668
-
Fastest Convergence for Q-learning 12 Jul 2017 · 0 repositories · arXiv:1707.03770
-
Noisy Networks for Exploration 30 Jun 2017 · 15 repositories · arXiv:1706.10295Syntology 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
A Self-Adaptive Proposal Model for Temporal Action Detection based on Reinforcement Learning 22 Jun 2017 · 1 repository · arXiv:1706.07251
-
Generalized Value Iteration Networks: Life Beyond Lattices 8 Jun 2017 · 1 repository · arXiv:1706.02416
-
Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments 7 Jun 2017 · 86 repositories · arXiv:1706.02275Syntology official: no sample here; runs from other or unrecorded repositories · 75 ran (of which 54 constructed an object rather than computing a result; 68 with no instrument failure: 2 honoured, 0 violated, 66 with no contract checked; 7 where Syntology's instrument failed) · 68 unverified (of 143 harvested samples) · 99 pointer-only (licence)
-
Parameter Space Noise for Exploration 6 Jun 2017 · 10 repositories · arXiv:1706.01905Syntology 4 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Explaining Transition Systems through Program Induction 23 May 2017 · 0 repositories · arXiv:1705.08320
-
Shallow Updates for Deep Reinforcement Learning 21 May 2017 · 0 repositories · arXiv:1705.07461
-
A Comparison of Reinforcement Learning Techniques for Fuzzy Cloud Auto-Scaling 19 May 2017 · 0 repositories · arXiv:1705.07114
-
Learning to Represent Haptic Feedback for Partially-Observable Tasks 17 May 2017 · 0 repositories · arXiv:1705.06243
-
Learning Hard Alignments with Variational Inference 16 May 2017 · 0 repositories · arXiv:1705.05524
-
Deep Episodic Value Iteration for Model-based Meta-Reinforcement Learning 9 May 2017 · 0 repositories · arXiv:1705.03562
-
Reinforcement Learning with External Knowledge and Two-Stage Q-functions for Predicting Popular Reddit Threads 20 Apr 2017 · 0 repositories · arXiv:1704.06217
-
The Reactor: A fast and sample-efficient Actor-Critic agent for Reinforcement Learning 15 Apr 2017 · 0 repositories · arXiv:1704.04651
-
Deep Q-learning from Demonstrations 12 Apr 2017 · 6 repositories · arXiv:1704.03732
-
Data-efficient Deep Reinforcement Learning for Dexterous Manipulation 10 Apr 2017 · 0 repositories · arXiv:1704.03073
-
Pseudorehearsal in value function approximation 21 Mar 2017 · 0 repositories · arXiv:1703.07075
-
Evolution Strategies as a Scalable Alternative to Reinforcement Learning 10 Mar 2017 · 23 repositories · arXiv:1703.03864Syntology official (archive's flag): 4 ran · 16 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 1 violated, 10 with no contract checked; 5 where Syntology's instrument failed) · 13 unverified (of 29 harvested samples) · 2 pointer-only (licence)
-
Tactics of Adversarial Attack on Deep Reinforcement Learning Agents 8 Mar 2017 · 0 repositories · arXiv:1703.06748
-
Count-Based Exploration with Neural Density Models 3 Mar 2017 · 1 repository · arXiv:1703.01310Syntology 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Bridging the Gap Between Value and Policy Based Reinforcement Learning 28 Feb 2017 · 1 repository · arXiv:1702.08892
-
Stabilising Experience Replay for Deep Multi-Agent Reinforcement Learning 28 Feb 2017 · 5 repositories · arXiv:1702.08887
-
Learning Control for Air Hockey Striking using Deep Reinforcement Learning 26 Feb 2017 · 0 repositories · arXiv:1702.08074
-
Sigmoid-Weighted Linear Units for Neural Network Function Approximation in Reinforcement Learning 10 Feb 2017 · 0 repositories · arXiv:1702.03118
-
Autonomous Braking System via Deep Reinforcement Learning 8 Feb 2017 · 2 repositories · arXiv:1702.02302
-
FPGA Architecture for Deep Learning and its application to Planetary Robotics 26 Jan 2017 · 0 repositories · arXiv:1701.07543
-
Vulnerability of Deep Reinforcement Learning to Policy Induction Attacks 16 Jan 2017 · 1 repository · arXiv:1701.04143
-
Deep Reinforcement Learning for Multi-Domain Dialogue Systems 26 Nov 2016 · 1 repository · arXiv:1611.08675
-
Memory Lens: How Much Memory Does an Agent Use? 21 Nov 2016 · 0 repositories · arXiv:1611.06928
-
Averaged-DQN: Variance Reduction and Stabilization for Deep Reinforcement Learning 7 Nov 2016 · 0 repositories · arXiv:1611.01929
-
Learning to Play in a Day: Faster Deep Reinforcement Learning by Optimality Tightening 5 Nov 2016 · 1 repository · arXiv:1611.01606
-
Combating Reinforcement Learning's Sisyphean Curse with Intrinsic Fear 3 Nov 2016 · 0 repositories · arXiv:1611.01211
-
Using a Deep Reinforcement Learning Agent for Traffic Signal Control 3 Nov 2016 · 0 repositories · arXiv:1611.01142
-
Internet of Things Applications: Animal Monitoring with Unmanned Aerial Vehicle 17 Oct 2016 · 0 repositories · arXiv:1610.05287
-
Deep Reinforcement Learning From Raw Pixels in Doom 7 Oct 2016 · 0 repositories · arXiv:1610.02164
-
Active exploration in parameterized reinforcement learning 6 Oct 2016 · 1 repository · arXiv:1610.01986
-
Opponent Modeling in Deep Reinforcement Learning 18 Sep 2016 · 1 repository · arXiv:1609.05559
-
Episodic Exploration for Deep Deterministic Policies: An Application to StarCraft Micromanagement Tasks 10 Sep 2016 · 0 repositories · arXiv:1609.02993
-
Multi Exit Configuration of Mesoscopic Pedestrian Simulation 6 Sep 2016 · 0 repositories · arXiv:1609.01475
-
BBQ-Networks: Efficient Exploration in Deep Reinforcement Learning for Task-Oriented Dialogue Systems 17 Aug 2016 · 0 repositories · arXiv:1608.05081
-
Deep Reinforcement Learning Discovers Internal Models 16 Jun 2016 · 0 repositories · arXiv:1606.05174
-
Deep Reinforcement Learning With Macro-Actions 15 Jun 2016 · 0 repositories · arXiv:1606.04615
-
ViZDoom: A Doom-based AI Research Platform for Visual Reinforcement Learning 6 May 2016 · 10 repositories · arXiv:1605.02097Syntology community repositories only · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Classifying Options for Deep Reinforcement Learning 27 Apr 2016 · 0 repositories · arXiv:1604.08153
-
Neurohex: A Deep Q-learning Hex Agent 24 Apr 2016 · 0 repositories · arXiv:1604.07097
-
Continuous Deep Q-Learning with Model-based Acceleration 2 Mar 2016 · 8 repositories · arXiv:1603.00748Syntology 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Reinforcement Learning approach for Real Time Strategy Games Battle city and S3 16 Feb 2016 · 0 repositories · arXiv:1602.04936