Methods › Reinforcement Learning › Off-Policy TD Control › Q-Learning › Papers, page 15
Q-Learning
Papers archive 2025-07-28
archive papers tagged: 1,734 · with a code link: 464 · where Syntology ran a sample: 126 (105 with a run with no instrument failure, 21 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (126 of 1,734 tagged: 105 with a run with no instrument failure, 21 where every run was a failure of Syntology's instrument)
Page 15 of 18: papers 1,401 to 1,500 of 1,734, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Multi Pseudo Q-learning Based Deterministic Policy Gradient for Tracking Control of Autonomous Underwater Vehicles 7 Sep 2019 · 0 repositories · arXiv:1909.03204
-
Reinforcement Learning for Joint Optimization of Multiple Rewards 6 Sep 2019 · 0 repositories · arXiv:1909.02940
-
Encoders and Decoders for Quantum Expander Codes Using Machine Learning 6 Sep 2019 · 0 repositories · arXiv:1909.02945
-
Q-DATA: Enhanced Traffic Flow Monitoring in Software-Defined Networks applying Q-learning 4 Sep 2019 · 0 repositories · arXiv:1909.01544
-
Solving Discounted Stochastic Two-Player Games with Near-Optimal Time and Sample Complexity 29 Aug 2019 · 0 repositories · arXiv:1908.11071
-
Intelligent Active Queue Management Using Explicit Congestion Notification 28 Aug 2019 · 1 repository · arXiv:1909.08386
-
Networked Control of Nonlinear Systems under Partial Observation Using Continuous Deep Q-Learning 28 Aug 2019 · 0 repositories · arXiv:1908.10722
-
STMARL: A Spatio-Temporal Multi-Agent Reinforcement Learning Approach for Cooperative Traffic Light Control 28 Aug 2019 · 0 repositories · arXiv:1908.10577
-
Performing Deep Recurrent Double Q-Learning for Atari Games 16 Aug 2019 · 2 repositories · arXiv:1908.06040
-
Learn How to Cook a New Recipe in a New House: Using Map Familiarization, Curriculum Learning, and Bandit Feedback to Learn Families of Text-Based Adventure Games 13 Aug 2019 · 1 repository · arXiv:1908.04777Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Large-Scale Traffic Signal Control Using a Novel Multi-Agent Reinforcement Learning 10 Aug 2019 · 0 repositories · arXiv:1908.03761
-
Q-MIND: Defeating Stealthy DoS Attacks in SDN with a Machine-learning based Defense Framework 27 Jul 2019 · 0 repositories · arXiv:1907.11887
-
An Optimistic Perspective on Offline Reinforcement Learning 10 Jul 2019 · 1 repository · arXiv:1907.04543
-
Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog 30 Jun 2019 · 1 repository · arXiv:1907.00456Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Q-Learning Inspired Self-Tuning for Energy Efficiency in HPC 26 Jun 2019 · 0 repositories · arXiv:1906.10970
-
Towards Empathic Deep Q-Learning 26 Jun 2019 · 1 repository · arXiv:1906.10918Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Learning Causal State Representations of Partially Observable Environments 25 Jun 2019 · 0 repositories · arXiv:1906.10437
-
In Hindsight: A Smooth Reward for Steady Exploration 24 Jun 2019 · 0 repositories · arXiv:1906.09781
-
Optimal Use of Experience in First Person Shooter Environments 24 Jun 2019 · 0 repositories · arXiv:1906.09734
-
Neural networks with motivation 23 Jun 2019 · 0 repositories · arXiv:1906.09528
-
Reinforcement Learning-Based Trajectory Design for the Aerial Base Stations 23 Jun 2019 · 0 repositories · arXiv:1906.09550
-
A Story of Two Streams: Reinforcement Learning Models from Human Behavior and Neuropsychiatry 21 Jun 2019 · 1 repository · arXiv:1906.11286
-
Split Q Learning: Reinforcement Learning with Two-Stream Rewards 21 Jun 2019 · 1 repository · arXiv:1906.12350
-
A Generalized Minimax Q-learning Algorithm for Two-Player Zero-Sum Stochastic Games 16 Jun 2019 · 0 repositories · arXiv:1906.06659
-
Boosting Soft Actor-Critic: Emphasizing Recent Experience without Forgetting the Past 10 Jun 2019 · 3 repositories · arXiv:1906.04009Syntology 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Escaping the State of Nature: A Hobbesian Approach to Cooperation in Multi-agent Reinforcement Learning 5 Jun 2019 · 0 repositories · arXiv:1906.09874
-
Exploration with Unreliable Intrinsic Reward in Multi-Agent Reinforcement Learning 5 Jun 2019 · 0 repositories · arXiv:1906.02138
-
Risk-Sensitive Compact Decision Trees for Autonomous Execution in Presence of Simulated Market Response 5 Jun 2019 · 0 repositories · arXiv:1906.02312
-
Reinforcement Learning with Low-Complexity Liquid State Machines 4 Jun 2019 · 1 repository · arXiv:1906.01695Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
Stabilizing Off-Policy Q-Learning via Bootstrapping Error Reduction 3 Jun 2019 · 3 repositories · arXiv:1906.00949Syntology 4 ran (of which 4 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified; every one of the 4 samples that ran constructed an object rather than computing a result (of 5 harvested samples) · 5 pointer-only (licence)
-
Analysis and Improvement of Adversarial Training in DQN Agents With Adversarially-Guided Exploration (AGE) 3 Jun 2019 · 0 repositories · arXiv:1906.01119
-
RL-Based Method for Benchmarking the Adversarial Resilience and Robustness of Deep Reinforcement Learning Policies 3 Jun 2019 · 0 repositories · arXiv:1906.01110
-
Sequential Triggers for Watermarking of Deep Reinforcement Learning Policies 3 Jun 2019 · 0 repositories · arXiv:1906.01126
-
Feature-Based Q-Learning for Two-Player Stochastic Games 2 Jun 2019 · 0 repositories · arXiv:1906.00423
-
Provably Efficient Q-Learning with Low Switching Cost 30 May 2019 · 0 repositories · arXiv:1905.12849
-
Reinforcement Learning for Slate-based Recommender Systems: A Tractable Decomposition and Practical Methodology 29 May 2019 · 3 repositories · arXiv:1905.12767
-
Learning distant cause and effect using only local and immediate credit assignment 28 May 2019 · 0 repositories · arXiv:1905.11589
-
Solving NP-Hard Problems on Graphs with Extended AlphaGo Zero 28 May 2019 · 2 repositories · arXiv:1905.11623Syntology 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Finite-Sample Analysis of Nonlinear Stochastic Approximation with Applications in Reinforcement Learning 27 May 2019 · 1 repository · arXiv:1905.11425Syntology 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
SQIL: Imitation Learning via Reinforcement Learning with Sparse Rewards 27 May 2019 · 5 repositories · arXiv:1905.11108
-
Prioritized Sequence Experience Replay 25 May 2019 · 0 repositories · arXiv:1905.12726
-
A Kernel Loss for Solving the Bellman Equation 25 May 2019 · 1 repository · arXiv:1905.10506
-
MQLV: Optimal Policy of Money Management in Retail Banking with Q-Learning 24 May 2019 · 0 repositories · arXiv:1905.12567
-
Adaptive Symmetric Reward Noising for Reinforcement Learning 24 May 2019 · 1 repository · arXiv:1905.10144
-
Deep Q-Learning with Q-Matrix Transfer Learning for Novel Fire Evacuation Environment 23 May 2019 · 0 repositories · arXiv:1905.09673
-
Deep Reinforcement Learning Based Parameter Control in Differential Evolution 20 May 2019 · 1 repository · arXiv:1905.08006
-
Stochastic Variance Reduction for Deep Q-learning 20 May 2019 · 0 repositories · arXiv:1905.08152
-
Reinforcement Learning for Learning of Dynamical Systems in Uncertain Environment: a Tutorial 19 May 2019 · 0 repositories · arXiv:1905.07727
-
Mastering the Game of Sungka from Random Play 17 May 2019 · 1 repository · arXiv:1905.07102
-
QBSO-FS: A Reinforcement Learning Based Bee Swarm Optimization Metaheuristic for Feature Selection 16 May 2019 · 1 repository
-
Autonomous Penetration Testing using Reinforcement Learning 15 May 2019 · 0 repositories · arXiv:1905.05965
-
Reinforcement Learning for Robotics and Control with Active Uncertainty Reduction 15 May 2019 · 0 repositories · arXiv:1905.06274
-
Design of Artificial Intelligence Agents for Games using Deep Reinforcement Learning 10 May 2019 · 0 repositories · arXiv:1905.04127
-
Domain Adversarial Reinforcement Learning for Partial Domain Adaptation 10 May 2019 · 0 repositories · arXiv:1905.04094
-
Pretrain Soft Q-Learning with Imperfect Demonstrations 9 May 2019 · 0 repositories · arXiv:1905.03501
-
A Reinforcement Learning Perspective on the Optimal Control of Mutation Probabilities for the (1+1) Evolutionary Algorithm: First Results on the OneMax Problem 9 May 2019 · 0 repositories · arXiv:1905.03726
-
Accelerated Target Updates for Q-learning 7 May 2019 · 0 repositories · arXiv:1905.02841
-
Comprehensible Context-driven Text Game Playing 6 May 2019 · 2 repositories · arXiv:1905.02265
-
Deep Ordinal Reinforcement Learning 6 May 2019 · 1 repository · arXiv:1905.02005
-
Beyond Games: Bringing Exploration to Robots in Real-world 1 May 2019 · 0 repositories
-
Efficient Model-free Reinforcement Learning in Metric Spaces 1 May 2019 · 1 repository · arXiv:1905.00475
-
Inducing Cooperation via Learning to reshape rewards in semi-cooperative multi-agent reinforcement learning 1 May 2019 · 0 repositories
-
Learning agents with prioritization and parameter noise in continuous state and action space 1 May 2019 · 0 repositories
-
Recurrent Experience Replay in Distributed Reinforcement Learning 1 May 2019 · 3 repositories
-
A Deep Q-Learning Method for Downlink Power Allocation in Multi-Cell Networks 30 Apr 2019 · 0 repositories · arXiv:1904.13032
-
Generative Adversarial Imagination for Sample Efficient Deep Reinforcement Learning 30 Apr 2019 · 0 repositories · arXiv:1904.13255
-
Zap Q-Learning for Optimal Stopping Time Problems 25 Apr 2019 · 0 repositories · arXiv:1904.11538
-
Stochastic Lipschitz Q-Learning 24 Apr 2019 · 0 repositories · arXiv:1904.10653
-
Target-Based Temporal Difference Learning 24 Apr 2019 · 0 repositories · arXiv:1904.10945
-
Deep Q Learning Driven CT Pancreas Segmentation with Geometry-Aware U-Net 19 Apr 2019 · 0 repositories · arXiv:1904.09120
-
"Jam Me If You Can'': Defeating Jammer with Deep Dueling Neural Network Architecture and Ambient Backscattering Augmented Communications 8 Apr 2019 · 0 repositories · arXiv:1904.03897
-
Personalized Cancer Chemotherapy Schedule: a numerical comparison of performance and robustness in model-based and model-free scheduling methodologies 2 Apr 2019 · 0 repositories · arXiv:1904.01200
-
Lane Change Decision-making through Deep Reinforcement Learning with Rule-based Constraints 30 Mar 2019 · 0 repositories · arXiv:1904.00231
-
Improved robustness of reinforcement learning policies upon conversion to spiking neuronal network platforms applied to ATARI games 26 Mar 2019 · 3 repositories · arXiv:1903.11012
-
Q-Learning for Continuous Actions with Cross-Entropy Guided Policies 25 Mar 2019 · 0 repositories · arXiv:1903.10605
-
DQN with model-based exploration: efficient learning on environments with sparse rewards 22 Mar 2019 · 0 repositories · arXiv:1903.09295
-
Towards Characterizing Divergence in Deep Q-Learning 21 Mar 2019 · 0 repositories · arXiv:1903.08894
-
Deep Reinforcement Learning with Decorrelation 18 Mar 2019 · 0 repositories · arXiv:1903.07765
-
Reinforcement Learning with Dynamic Boltzmann Softmax Updates 14 Mar 2019 · 1 repository · arXiv:1903.05926
-
Deep Multi-Agent Reinforcement Learning with Discrete-Continuous Hybrid Action Spaces 12 Mar 2019 · 0 repositories · arXiv:1903.04959
-
Deep Recurrent Q-Learning vs Deep Q-Learning on a simple Partially Observable Markov Decision Process with Minecraft 11 Mar 2019 · 2 repositories · arXiv:1903.04311
-
Multi-Agent Deep Reinforcement Learning for Large-scale Traffic Signal Control 11 Mar 2019 · 1 repository · arXiv:1903.04527Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 9 harvested samples) · 1 pointer-only (licence)
-
Sample-Efficient Model-Free Reinforcement Learning with Off-Policy Critics 11 Mar 2019 · 1 repository · arXiv:1903.04193
-
DeepPool: Distributed Model-free Algorithm for Ride-sharing using Deep Reinforcement Learning 9 Mar 2019 · 0 repositories · arXiv:1903.03882
-
Successive Over Relaxation Q-Learning 9 Mar 2019 · 0 repositories · arXiv:1903.03812
-
Learning Heuristics over Large Graphs via Deep Reinforcement Learning 8 Mar 2019 · 2 repositories · arXiv:1903.03332
-
MinAtar: An Atari-Inspired Testbed for Thorough and Reproducible Reinforcement Learning Experiments 7 Mar 2019 · 3 repositories · arXiv:1903.03176Syntology official: harvested, nothing ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Distributed Edge Caching via Reinforcement Learning in Fog Radio Access Networks 27 Feb 2019 · 0 repositories · arXiv:1902.10574
-
Unifying Ensemble Methods for Q-learning via Social Choice Theory 27 Feb 2019 · 0 repositories · arXiv:1902.10646
-
Diagnosing Bottlenecks in Deep Q-learning Algorithms 26 Feb 2019 · 1 repository · arXiv:1902.10250Syntology 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 10 harvested samples)
-
Optimal and Fast Real-time Resources Slicing with Deep Dueling Neural Networks 26 Feb 2019 · 0 repositories · arXiv:1902.09696
-
Autonomous Airline Revenue Management: A Deep Reinforcement Learning Approach to Seat Inventory Control and Overbooking 18 Feb 2019 · 0 repositories · arXiv:1902.06824
-
Heuristics, Answer Set Programming and Markov Decision Process for Solving a Set of Spatial Puzzles 16 Feb 2019 · 1 repository · arXiv:1903.03411
-
Sample-Optimal Parametric Q-Learning Using Linearly Additive Features 13 Feb 2019 · 0 repositories · arXiv:1902.04779
-
Learning Best Response Strategies for Agents in Ad Exchanges 10 Feb 2019 · 0 repositories · arXiv:1902.03588
-
Making Deep Q-learning methods robust to time discretization 28 Jan 2019 · 1 repository · arXiv:1901.09732Syntology 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Q-learning with UCB Exploration is Sample Efficient for Infinite-Horizon MDP 27 Jan 2019 · 0 repositories · arXiv:1901.09311
-
Reward Shaping via Meta-Learning 27 Jan 2019 · 0 repositories · arXiv:1901.09330
-
Combinational Q-Learning for Dou Di Zhu 24 Jan 2019 · 1 repository · arXiv:1901.08925
-
Distillation Strategies for Proximal Policy Optimization 23 Jan 2019 · 0 repositories · arXiv:1901.08128