Methods › Reinforcement Learning › Off-Policy TD Control › Q-Learning › Papers, page 13
Q-Learning
Papers archive 2025-07-28
archive papers tagged: 1,734 · with a code link: 464 · where Syntology ran a sample: 126 (105 with a run with no instrument failure, 21 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (126 of 1,734 tagged: 105 with a run with no instrument failure, 21 where every run was a failure of Syntology's instrument)
Page 13 of 18: papers 1,201 to 1,300 of 1,734, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Reward Machines for Cooperative Multi-Agent Reinforcement Learning 3 Jul 2020 · 2 repositories · arXiv:2007.01962
-
Decentralized Deep Reinforcement Learning for Network Level Traffic Signal Control 2 Jul 2020 · 0 repositories · arXiv:2007.03433
-
Gradient Temporal-Difference Learning with Regularized Corrections 1 Jul 2020 · 1 repository · arXiv:2007.00611Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Group Equivariant Deep Reinforcement Learning 1 Jul 2020 · 1 repository · arXiv:2007.03437Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Regularly Updated Deterministic Policy Gradient Algorithm 1 Jul 2020 · 0 repositories · arXiv:2007.00169
-
Provably More Efficient Q-Learning in the One-Sided-Feedback/Full-Feedback Settings 30 Jun 2020 · 0 repositories · arXiv:2007.00080
-
Concept and the implementation of a tool to convert industry 4.0 environments modeled as FSM to an OpenAI Gym wrapper 29 Jun 2020 · 0 repositories · arXiv:2006.16035
-
Using Reinforcement Learning to Herd a Robotic Swarm to a Target Distribution 29 Jun 2020 · 0 repositories · arXiv:2006.15807
-
Image Classification by Reinforcement Learning with Two-State Q-Learning 28 Jun 2020 · 1 repository · arXiv:2007.01298
-
Lookahead-Bounded Q-Learning 28 Jun 2020 · 1 repository · arXiv:2006.15690
-
Reinforcement Learning Based Handwritten Digit Recognition with Two-State Q-Learning 28 Jun 2020 · 0 repositories · arXiv:2007.01193
-
Offline Contextual Bandits with Overparameterized Models 27 Jun 2020 · 1 repository · arXiv:2006.15368
-
Q-Learning with Differential Entropy of Q-Tables 26 Jun 2020 · 0 repositories · arXiv:2006.14795
-
Some approaches used to overcome overestimation in Deep Reinforcement Learning algorithms 25 Jun 2020 · 0 repositories · arXiv:2006.14167
-
Preventing Value Function Collapse in Ensemble Q-Learning by Maximizing Representation Diversity 24 Jun 2020 · 0 repositories · arXiv:2006.13823
-
RL Unplugged: A Suite of Benchmarks for Offline Reinforcement Learning 24 Jun 2020 · 2 repositories · arXiv:2006.13888
-
Deep Reinforcement Learning Control for Radar Detection and Tracking in Congested Spectral Environments 23 Jun 2020 · 0 repositories · arXiv:2006.13173
-
Risk-Sensitive Reinforcement Learning: Near-Optimal Risk-Sample Tradeoff in Regret 22 Jun 2020 · 0 repositories · arXiv:2006.13827
-
NROWAN-DQN: A Stable Noisy Network with Noise Reduction and Online Weight Adjustment for Exploration 19 Jun 2020 · 0 repositories · arXiv:2006.10980
-
Efficient Ridesharing Dispatch Using Multi-Agent Reinforcement Learning 18 Jun 2020 · 1 repository · arXiv:2006.10897
-
Parameterized MDPs and Reinforcement Learning Problems -- A Maximum Entropy Principle Based Framework 17 Jun 2020 · 0 repositories · arXiv:2006.09646
-
Semantic Visual Navigation by Watching YouTube Videos 17 Jun 2020 · 1 repository · arXiv:2006.10034
-
The Sample Complexity of Teaching-by-Reinforcement on Q-Learning 16 Jun 2020 · 0 repositories · arXiv:2006.09324
-
Interaction Networks: Using a Reinforcement Learner to train other Machine Learning algorithms 15 Jun 2020 · 0 repositories · arXiv:2006.08457
-
Decorrelated Double Q-learning 12 Jun 2020 · 0 repositories · arXiv:2006.06956
-
Deep Reinforcement Learning for Neural Control 12 Jun 2020 · 0 repositories · arXiv:2006.07352
-
Human and Multi-Agent collaboration in a human-MARL teaming framework 12 Jun 2020 · 0 repositories · arXiv:2006.07301
-
Safety-guaranteed Reinforcement Learning based on Multi-class Support Vector Machine 12 Jun 2020 · 0 repositories · arXiv:2006.07446
-
Self-Imitation Learning via Generalized Lower Bound Q-learning 12 Jun 2020 · 0 repositories · arXiv:2006.07442Syntology 16 ran (of which 4 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 2 violated, 11 with no contract checked; 3 where Syntology's instrument failed) · 4 unverified (of 20 harvested samples) · 11 pointer-only (licence)
-
Exploration by Maximizing Rényi Entropy for Reward-Free RL Framework 11 Jun 2020 · 0 repositories · arXiv:2006.06193
-
Fitted Q-Learning for Relational Domains 10 Jun 2020 · 0 repositories · arXiv:2006.05595
-
Model-Free Algorithm and Regret Analysis for MDPs with Long-Term Constraints 10 Jun 2020 · 0 repositories · arXiv:2006.05961
-
Multi-Agent Reinforcement Learning in a Realistic Limit Order Book Market Simulation 10 Jun 2020 · 0 repositories · arXiv:2006.05574
-
Privacy-Cost Management in Smart Meters with Mutual Information-Based Reinforcement Learning 10 Jun 2020 · 0 repositories · arXiv:2006.06106
-
Reinforcement Learning-Based Joint Self-Optimisation Method for the Fuzzy Logic Handover Algorithm in 5G HetNets 9 Jun 2020 · 0 repositories · arXiv:2006.05010
-
A Model-free Learning Algorithm for Infinite-horizon Average-reward MDPs with Near-optimal Regret 8 Jun 2020 · 0 repositories · arXiv:2006.04354
-
Balancing a CartPole System with Reinforcement Learning -- A Tutorial 8 Jun 2020 · 0 repositories · arXiv:2006.04938
-
Can Temporal-Difference and Q-Learning Learn Representation? A Mean-Field Theory 8 Jun 2020 · 0 repositories · arXiv:2006.04761
-
Conservative Q-Learning for Offline Reinforcement Learning 8 Jun 2020 · 18 repositories · arXiv:2006.04779Syntology community repositories only · 31 ran (of which 5 constructed an object rather than computing a result; 28 with no instrument failure: 2 honoured, 0 violated, 26 with no contract checked; 3 where Syntology's instrument failed) · 3 unverified (of 34 harvested samples) · 19 pointer-only (licence)
-
A Multi-step and Resilient Predictive Q-learning Algorithm for IoT with Human Operators in the Loop: A Case Study in Water Supply Networks 6 Jun 2020 · 0 repositories · arXiv:2006.03899
-
Logical Team Q-learning: An approach towards factored policies in cooperative MARL 5 Jun 2020 · 0 repositories · arXiv:2006.03553
-
A Novel Update Mechanism for Q-Networks Based On Extreme Learning Machines 4 Jun 2020 · 1 repository · arXiv:2006.02986
-
Sample Complexity of Asynchronous Q-Learning: Sharper Analysis and Variance Reduction 4 Jun 2020 · 0 repositories · arXiv:2006.03041
-
Mitigating Bias in Face Recognition Using Skewness-Aware Reinforcement Learning 1 Jun 2020 · 0 repositories
-
Towards Understanding Cooperative Multi-Agent Q-Learning with Value Factorization 31 May 2020 · 0 repositories · arXiv:2006.00587
-
Active Measure Reinforcement Learning for Observation Cost Minimization 26 May 2020 · 0 repositories · arXiv:2005.12697
-
Learning to Charge RF-Energy Harvesting Devices in WiFi Networks 25 May 2020 · 0 repositories · arXiv:2005.12022
-
Should artificial agents ask for help in human-robot collaborative problem-solving? 25 May 2020 · 0 repositories · arXiv:2006.00882
-
A reinforcement learning based decision support system in textile manufacturing process 20 May 2020 · 0 repositories · arXiv:2005.09867
-
Prototypical Q Networks for Automatic Conversational Diagnosis and Few-Shot New Disease Adaption 19 May 2020 · 0 repositories · arXiv:2005.11153
-
Basal Glucose Control in Type 1 Diabetes using Deep Reinforcement Learning: An In Silico Validation 18 May 2020 · 0 repositories · arXiv:2005.09059
-
Local and Global Explanations of Agent Behavior: Integrating Strategy Summaries with Saliency Maps 18 May 2020 · 1 repository · arXiv:2005.08874Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 6 unverified (of 15 harvested samples)
-
A Deep Q-learning/genetic Algorithms Based Novel Methodology For Optimizing Covid-19 Pandemic Government Actions 15 May 2020 · 0 repositories · arXiv:2005.07656
-
A Deep Reinforcement Learning Approach to Efficient Drone Mobility Support 11 May 2020 · 0 repositories · arXiv:2005.05229
-
An FPGA-Based On-Device Reinforcement Learning Approach using Online Sequential Learning 10 May 2020 · 0 repositories · arXiv:2005.04646
-
Reinforcement Learning for Thermostatically Controlled Loads Control using Modelica and Python 9 May 2020 · 0 repositories · arXiv:2005.04444
-
Optimal Beam Association for High Mobility mmWave Vehicular Networks: Lightweight Parallel Reinforcement Learning Approach 2 May 2020 · 0 repositories · arXiv:2005.00694
-
Implementing Inductive bias for different navigation tasks through diverse RNN attrractors 1 May 2020 · 0 repositories
-
Learning Efficient Parameter Server Synchronization Policies for Distributed SGD 1 May 2020 · 0 repositories
-
Whittle index based Q-learning for restless bandits with average reward 29 Apr 2020 · 0 repositories · arXiv:2004.14427
-
Learning Dialog Policies from Weak Demonstrations 23 Apr 2020 · 0 repositories · arXiv:2004.11054
-
Spatial Action Maps for Mobile Manipulation 20 Apr 2020 · 1 repository · arXiv:2004.09141Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 8 harvested samples) · 1 pointer-only (licence)
-
Deep Reinforcement Learning for Adaptive Learning Systems 17 Apr 2020 · 0 repositories · arXiv:2004.08410
-
Show Us the Way: Learning to Manage Dialog from Demonstrations 17 Apr 2020 · 0 repositories · arXiv:2004.08114
-
K-spin Hamiltonian for quantum-resolvable Markov decision processes 13 Apr 2020 · 0 repositories · arXiv:2004.06040
-
Risk-Aware High-level Decisions for Automated Driving at Occluded Intersections with Reinforcement Learning 9 Apr 2020 · 0 repositories · arXiv:2004.04450
-
An Application of Deep Reinforcement Learning to Algorithmic Trading 7 Apr 2020 · 1 repository · arXiv:2004.06627
-
Uniform State Abstraction For Reinforcement Learning 6 Apr 2020 · 0 repositories · arXiv:2004.02919
-
Zero-Shot Learning of Text Adventure Games with Sentence-Level Semantics 6 Apr 2020 · 0 repositories · arXiv:2004.02986
-
Multi-agent Reinforcement Learning for Resource Allocation in IoT networks with Edge Computing 5 Apr 2020 · 0 repositories · arXiv:2004.02315
-
Minimizing Age-of-Information for Fog Computing-supported Vehicular Networks with Deep Q-learning 4 Apr 2020 · 0 repositories · arXiv:2004.04640
-
Reinforcement Learning for Mixed-Integer Problems Based on MPC 3 Apr 2020 · 0 repositories · arXiv:2004.01430
-
Statistically Model Checking PCTL Specifications on Markov Decision Processes via Reinforcement Learning 1 Apr 2020 · 0 repositories · arXiv:2004.00273
-
Enhanced Rolling Horizon Evolution Algorithm with Opponent Model Learning: Results for the Fighting Game AI Competition 31 Mar 2020 · 0 repositories · arXiv:2003.13949
-
Robust Q-learning 27 Mar 2020 · 0 repositories · arXiv:2003.12427
-
Importance of using appropriate baselines for evaluation of data-efficiency in deep reinforcement learning for Atari 23 Mar 2020 · 0 repositories · arXiv:2003.10181
-
Using Deep Reinforcement Learning Methods for Autonomous Vessels in 2D Environments 23 Mar 2020 · 1 repository · arXiv:2003.10249
-
Distributed Reinforcement Learning for Cooperative Multi-Robot Object Manipulation 21 Mar 2020 · 0 repositories · arXiv:2003.09540
-
FlapAI Bird: Training an Agent to Play Flappy Bird Using Reinforcement Learning Techniques 21 Mar 2020 · 2 repositories · arXiv:2003.09579
-
Deep Reinforcement Learning with Weighted Q-Learning 20 Mar 2020 · 0 repositories · arXiv:2003.09280
-
Deep Constrained Q-learning 20 Mar 2020 · 0 repositories · arXiv:2003.09398
-
Robust Deep Reinforcement Learning against Adversarial Perturbations on State Observations 19 Mar 2020 · 4 repositories · arXiv:2003.08938
-
Simultaneous Navigation and Radio Mapping for Cellular-Connected UAV with Deep Reinforcement Learning 17 Mar 2020 · 1 repository · arXiv:2003.07574
-
DisCor: Corrective Feedback in Reinforcement Learning via Distribution Correction 16 Mar 2020 · 4 repositories · arXiv:2003.07305Syntology 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
Active Perception and Representation for Robotic Manipulation 15 Mar 2020 · 0 repositories · arXiv:2003.06734
-
A General Framework for Learning Mean-Field Games 13 Mar 2020 · 0 repositories · arXiv:2003.06069
-
Application of Deep Q-Network in Portfolio Management 13 Mar 2020 · 0 repositories · arXiv:2003.06365
-
Provably Efficient Model-Free Algorithm for MDPs with Peak Constraints 11 Mar 2020 · 0 repositories · arXiv:2003.05555
-
A Multi-Agent Reinforcement Learning Approach For Safe and Efficient Behavior Planning Of Connected Autonomous Vehicles 9 Mar 2020 · 0 repositories · arXiv:2003.04371
-
Software-Level Accuracy Using Stochastic Computing With Charge-Trap-Flash Based Weight Matrix 9 Mar 2020 · 0 repositories · arXiv:2004.11120
-
Dynamic Experience Replay 4 Mar 2020 · 0 repositories · arXiv:2003.02372
-
Contention Window Optimization in IEEE 802.11ax Networks with Deep Reinforcement Learning 3 Mar 2020 · 1 repository · arXiv:2003.01492
-
Relevance-Guided Modeling of Object Dynamics for Reinforcement Learning 3 Mar 2020 · 0 repositories · arXiv:2003.01384
-
Deep Reinforcement Learning for FlipIt Security Game 28 Feb 2020 · 0 repositories · arXiv:2002.12909
-
ConQUR: Mitigating Delusional Bias in Deep Q-learning 27 Feb 2020 · 1 repository · arXiv:2002.12399
-
Optimistic Exploration even with a Pessimistic Initialisation 26 Feb 2020 · 1 repository · arXiv:2002.12174Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
G-Learner and GIRL: Goal Based Wealth Management with Reinforcement Learning 25 Feb 2020 · 0 repositories · arXiv:2002.10990
-
A Double Q-Learning Approach for Navigation of Aerial Vehicles with Connectivity Constraint 24 Feb 2020 · 0 repositories · arXiv:2002.10563
-
Millimeter Wave Communications with an Intelligent Reflector: Performance Optimization and Distributional Reinforcement Learning 24 Feb 2020 · 0 repositories · arXiv:2002.10572
-
Q-learning with Uniformly Bounded Variance: Large Discounting is Not a Barrier to Fast Learning 24 Feb 2020 · 0 repositories · arXiv:2002.10301