Methods › Reinforcement Learning › Off-Policy TD Control › Q-Learning › Papers, page 12
Q-Learning
Papers archive 2025-07-28
archive papers tagged: 1,734 · with a code link: 464 · where Syntology ran a sample: 126 (105 with a run with no instrument failure, 21 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (126 of 1,734 tagged: 105 with a run with no instrument failure, 21 where every run was a failure of Syntology's instrument)
Page 12 of 18: papers 1,101 to 1,200 of 1,734, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Real-time Active Vision for a Humanoid Soccer Robot Using Deep Reinforcement Learning 27 Nov 2020 · 0 repositories · arXiv:2011.13851
-
Reinforcement Learning-based Joint Path and Energy Optimization of Cellular-Connected Unmanned Aerial Vehicles 27 Nov 2020 · 0 repositories · arXiv:2011.13744
-
Predictive PER: Balancing Priority and Diversity towards Stable Deep Reinforcement Learning 26 Nov 2020 · 0 repositories · arXiv:2011.13093
-
Diluted Near-Optimal Expert Demonstrations for Guiding Dialogue Stochastic Policy Optimisation 25 Nov 2020 · 0 repositories · arXiv:2012.04687
-
Learning Principle of Least Action with Reinforcement Learning 24 Nov 2020 · 1 repository · arXiv:2011.11891
-
Solving The Lunar Lander Problem under Uncertainty using Reinforcement Learning 24 Nov 2020 · 2 repositories · arXiv:2011.11850
-
Revisiting Rainbow: Promoting more Insightful and Inclusive Deep Reinforcement Learning Research 20 Nov 2020 · 2 repositories · arXiv:2011.14826
-
Adaptive Contention Window Design using Deep Q-learning 18 Nov 2020 · 1 repository · arXiv:2011.09418
-
Leveraging the Variance of Return Sequences for Exploration Policy 17 Nov 2020 · 0 repositories · arXiv:2011.08649
-
Constrained Model-Free Reinforcement Learning for Process Optimization 16 Nov 2020 · 0 repositories · arXiv:2011.07925
-
A deep Q-Learning based Path Planning and Navigation System for Firefighting Environments 12 Nov 2020 · 0 repositories · arXiv:2011.06450
-
Optimizing Large-Scale Fleet Management on a Road Network using Multi-Agent Deep Reinforcement Learning with Graph Neural Network 12 Nov 2020 · 1 repository · arXiv:2011.06175
-
On Using Hamiltonian Monte Carlo Sampling for Reinforcement Learning Problems in High-dimension 11 Nov 2020 · 0 repositories · arXiv:2011.05927
-
Multi-Agent Reinforcement Learning for Channel Assignment and Power Allocation in Platoon-Based C-V2X Systems 9 Nov 2020 · 0 repositories · arXiv:2011.04555
-
Reinforced Deep Markov Models With Applications in Automatic Trading 9 Nov 2020 · 0 repositories · arXiv:2011.04391
-
Reinforcement Learning for Assignment problem 8 Nov 2020 · 0 repositories · arXiv:2011.03909
-
Control with adaptive Q-learning 3 Nov 2020 · 1 repository · arXiv:2011.02141
-
DeepFoldit -- A Deep Reinforcement Learning Neural Network Folding Proteins 28 Oct 2020 · 0 repositories · arXiv:2011.03442
-
Finite-Time Convergence Rates of Decentralized Stochastic Approximation with Applications in Multi-Agent and Multi-Task Learning 28 Oct 2020 · 0 repositories · arXiv:2010.15088
-
Energy Consumption and Battery Aging Minimization Using a Q-learning Strategy for a Battery/Ultracapacitor Electric Vehicle 27 Oct 2020 · 0 repositories · arXiv:2010.14115
-
Hamilton-Jacobi Deep Q-Learning for Deterministic Continuous-Time Systems with Lipschitz Continuous Controls 27 Oct 2020 · 1 repository · arXiv:2010.14087
-
Learning Time Reduction Using Warm Start Methods for a Reinforcement Learning Based Supervisory Control in Hybrid Electric Vehicle Applications 27 Oct 2020 · 0 repositories · arXiv:2010.14575
-
Energy and Service-priority aware Trajectory Design for UAV-BSs using Double Q-Learning 26 Oct 2020 · 0 repositories · arXiv:2010.13346
-
Enhancing reinforcement learning by a finite reward response filter with a case study in intelligent structural control 25 Oct 2020 · 0 repositories · arXiv:2010.15597
-
Learning Guidance Rewards with Trajectory-space Smoothing 23 Oct 2020 · 2 repositories · arXiv:2010.12718Syntology community repositories only · 5 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Stabilizing Transformer-Based Action Sequence Generation For Q-Learning 23 Oct 2020 · 0 repositories · arXiv:2010.12698
-
Adversarial Attacks on Deep Algorithmic Trading Policies 22 Oct 2020 · 0 repositories · arXiv:2010.11388
-
A Weighted Heterogeneous Graph Based Dialogue System 21 Oct 2020 · 0 repositories · arXiv:2010.10699
-
Deep Surrogate Q-Learning for Autonomous Driving 21 Oct 2020 · 0 repositories · arXiv:2010.11278
-
On Information Asymmetry in Competitive Multi-Agent Reinforcement Learning: Convergence and Optimality 21 Oct 2020 · 0 repositories · arXiv:2010.10901
-
Language Inference with Multi-head Automata through Reinforcement Learning 20 Oct 2020 · 0 repositories · arXiv:2010.10141
-
Chance-Constrained Control with Lexicographic Deep Reinforcement Learning 19 Oct 2020 · 0 repositories · arXiv:2010.09468
-
Connections between Relational Event Model and Inverse Reinforcement Learning for Characterizing Group Interaction Sequences 19 Oct 2020 · 1 repository · arXiv:2010.09810
-
Multi-Agent Reinforcement Learning in NOMA-aided UAV Networks for Cellular Offloading 18 Oct 2020 · 0 repositories · arXiv:2010.09094
-
NOMA in UAV-aided cellular offloading: A machine learning approach 18 Oct 2020 · 0 repositories · arXiv:2011.14776
-
Learning Dexterous Manipulation from Suboptimal Experts 16 Oct 2020 · 0 repositories · arXiv:2010.08587
-
Multi-Agent Collaboration via Reward Attribution Decomposition 16 Oct 2020 · 2 repositories · arXiv:2010.08531
-
A Nesterov's Accelerated quasi-Newton method for Global Routing using Deep Reinforcement Learning 15 Oct 2020 · 0 repositories · arXiv:2010.09465
-
EpidemiOptim: A Toolbox for the Optimization of Control Policies in Epidemiological Models 9 Oct 2020 · 2 repositories · arXiv:2010.04452
-
Instance Weighted Incremental Evolution Strategies for Reinforcement Learning in Dynamic Environments 9 Oct 2020 · 1 repository · arXiv:2010.04605
-
Parameterized Reinforcement Learning for Optical System Optimization 9 Oct 2020 · 0 repositories · arXiv:2010.05769
-
Q-learning with Language Model for Edit-based Unsupervised Summarization 9 Oct 2020 · 1 repository · arXiv:2010.04379Syntology official: no sample here; runs from other or unrecorded repositories · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Model-Free Non-Stationary RL: Near-Optimal Regret and Applications in Multi-Agent RL and Inventory Control 7 Oct 2020 · 0 repositories · arXiv:2010.03161
-
Machine Learning Empowered Trajectory and Passive Beamforming Design in UAV-RIS Wireless Networks 6 Oct 2020 · 0 repositories · arXiv:2010.02749
-
Bayesian Meta-reinforcement Learning for Traffic Signal Control 1 Oct 2020 · 0 repositories · arXiv:2010.00163
-
Strategy and Benchmark for Converting Deep Q-Networks to Event-Driven Spiking Neural Networks 30 Sep 2020 · 0 repositories · arXiv:2009.14456
-
Cross Learning in Deep Q-Networks 29 Sep 2020 · 0 repositories · arXiv:2009.13780
-
Finite-Time Analysis for Double Q-learning 29 Sep 2020 · 0 repositories · arXiv:2009.14257
-
Lineage Evolution Reinforcement Learning 26 Sep 2020 · 0 repositories · arXiv:2010.14616
-
A New Approach for Tactical Decision Making in Lane Changing: Sample Efficient Deep Q Learning with a Safety Feedback Reward 24 Sep 2020 · 0 repositories · arXiv:2009.11905
-
Is Q-Learning Provably Efficient? An Extended Analysis 22 Sep 2020 · 0 repositories · arXiv:2009.10396
-
Hidden Incentives for Auto-Induced Distributional Shift 19 Sep 2020 · 0 repositories · arXiv:2009.09153
-
Energy-based Surprise Minimization for Multi-Agent Value Factorization 16 Sep 2020 · 1 repository · arXiv:2009.09842
-
Reinforcement Learning for Dynamic Resource Optimization in 5G Radio Access Network Slicing 14 Sep 2020 · 0 repositories · arXiv:2009.06579
-
AoI Minimization in Status Update Control with Energy Harvesting Sensors 9 Sep 2020 · 0 repositories · arXiv:2009.04224
-
Tactical Decision Making for Emergency Vehicles Based on A Combinational Learning Method 9 Sep 2020 · 0 repositories · arXiv:2009.04203
-
A Hybrid PAC Reinforcement Learning Algorithm 5 Sep 2020 · 0 repositories · arXiv:2009.02602
-
PAC Reinforcement Learning Algorithm for General-Sum Markov Games 5 Sep 2020 · 0 repositories · arXiv:2009.02605
-
DRLE: Decentralized Reinforcement Learning at the Edge for Traffic Light Control in the IoV 3 Sep 2020 · 1 repository · arXiv:2009.01502
-
Learning Nash Equilibria in Zero-Sum Stochastic Games via Entropy-Regularized Policy Approximation 1 Sep 2020 · 0 repositories · arXiv:2009.00162
-
Solving the single-track train scheduling problem via Deep Reinforcement Learning 1 Sep 2020 · 0 repositories · arXiv:2009.00433
-
Deep Q-Learning: Theoretical Insights from an Asymptotic Analysis 25 Aug 2020 · 0 repositories · arXiv:2008.10870
-
Table2Charts: Recommending Charts by Learning Shared Table Representations 24 Aug 2020 · 1 repository · arXiv:2008.11015
-
An adaptive synchronization approach for weights of deep reinforcement learning 16 Aug 2020 · 0 repositories · arXiv:2008.06973
-
The reinforcement learning-based multi-agent cooperative approach for the adaptive speed regulation on a metallurgical pickling line 16 Aug 2020 · 0 repositories · arXiv:2008.06933
-
Chrome Dino Run using Reinforcement Learning 15 Aug 2020 · 0 repositories · arXiv:2008.06799
-
Reinforcement Learning with Quantum Variational Circuits 15 Aug 2020 · 2 repositories · arXiv:2008.07524
-
Decision-making at Unsignalized Intersection for Autonomous Vehicles: Left-turn Maneuver with Deep Reinforcement Learning 14 Aug 2020 · 0 repositories · arXiv:2008.06595
-
Generalized Radio Environment Monitoring for Next Generation Wireless Networks 14 Aug 2020 · 0 repositories · arXiv:2008.06203
-
Multi-Agent Double Deep Q-Learning for Beamforming in mmWave MIMO Networks 13 Aug 2020 · 0 repositories · arXiv:2008.05943
-
Caching Placement and Resource Allocation for Cache-Enabling UAV NOMA Networks 12 Aug 2020 · 0 repositories · arXiv:2008.05168
-
Convex Q-Learning, Part 1: Deterministic Optimal Control 8 Aug 2020 · 0 repositories · arXiv:2008.03559
-
Evaluating Load Models and Their Impacts on Power Transfer Limits 7 Aug 2020 · 0 repositories · arXiv:2008.03336
-
Deep Q-Network Based Multi-agent Reinforcement Learning with Binary Action Agents 6 Aug 2020 · 0 repositories · arXiv:2008.04109
-
Deep Inverse Q-learning with Constraints 4 Aug 2020 · 2 repositories · arXiv:2008.01712
-
Cooperative Control of Mobile Robots with Stackelberg Learning 3 Aug 2020 · 0 repositories · arXiv:2008.00679
-
QPLEX: Duplex Dueling Multi-Agent Q-Learning 3 Aug 2020 · 6 repositories · arXiv:2008.01062Syntology community repositories only · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 1 pointer-only (licence)
-
Momentum Q-learning with Finite-Sample Convergence Guarantee 30 Jul 2020 · 0 repositories · arXiv:2007.15418
-
Deep Reinforcement Learning for Dynamic Spectrum Sensing and Aggregation in Multi-Channel Wireless Networks 28 Jul 2020 · 0 repositories · arXiv:2007.13965
-
Variance Reduction for Deep Q-Learning using Stochastic Recursive Gradient 25 Jul 2020 · 0 repositories · arXiv:2007.12817
-
A Comparative Study of AI-based Intrusion Detection Techniques in Critical Infrastructures 24 Jul 2020 · 0 repositories · arXiv:2008.00088
-
Trade-off on Sim2Real Learning: Real-world Learning Faster than Simulations 21 Jul 2020 · 0 repositories · arXiv:2007.10675
-
EMaQ: Expected-Max Q-Learning Operator for Simple Yet Effective Offline and Online RL 21 Jul 2020 · 0 repositories · arXiv:2007.11091
-
A Machine Learning Approach for Task and Resource Allocation in Mobile Edge Computing Based Networks 20 Jul 2020 · 0 repositories · arXiv:2007.10102
-
Multi-agent Reinforcement Learning in Bayesian Stackelberg Markov Games for Adaptive Moving Target Defense 20 Jul 2020 · 0 repositories · arXiv:2007.10457
-
DRIFT: Deep Reinforcement Learning for Functional Software Testing 16 Jul 2020 · 0 repositories · arXiv:2007.08220
-
Meta-Gradient Reinforcement Learning with an Objective Discovered Online 16 Jul 2020 · 0 repositories · arXiv:2007.08433
-
Mixture of Step Returns in Bootstrapped DQN 16 Jul 2020 · 0 repositories · arXiv:2007.08229
-
PC-PG: Policy Cover Directed Exploration for Provable Policy Gradient Learning 16 Jul 2020 · 1 repository · arXiv:2007.08459Syntology 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 1 honoured, 0 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 9 harvested samples) · 6 pointer-only (licence)
-
Reinforcement Learning-Enabled Decision-Making Strategies for a Vehicle-Cyber-Physical-System in Connected Environment 16 Jul 2020 · 0 repositories · arXiv:2007.09101
-
Analysis of Q-learning with Adaptation and Momentum Restart for Gradient Descent 15 Jul 2020 · 0 repositories · arXiv:2007.07422
-
Qgraph-bounded Q-learning: Stabilizing Model-Free Off-Policy Deep Reinforcement Learning 15 Jul 2020 · 0 repositories · arXiv:2007.07582
-
Single-partition adaptive Q-learning 14 Jul 2020 · 1 repository · arXiv:2007.06741
-
Revisiting Fundamentals of Experience Replay 13 Jul 2020 · 2 repositories · arXiv:2007.06700Syntology community repositories only · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Simulating multi-exit evacuation using deep reinforcement learning 11 Jul 2020 · 0 repositories · arXiv:2007.05783
-
The Mean-Squared Error of Double Q-Learning 9 Jul 2020 · 1 repository · arXiv:2007.05034Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
SUNRISE: A Simple Unified Framework for Ensemble Learning in Deep Reinforcement Learning 9 Jul 2020 · 1 repository · arXiv:2007.04938
-
Auto-MAP: A DQN Framework for Exploring Distributed Execution Plans for DNN Workloads 8 Jul 2020 · 0 repositories · arXiv:2007.04069
-
Cognitive Radio Network Throughput Maximization with Deep Reinforcement Learning 7 Jul 2020 · 0 repositories · arXiv:2007.03165
-
Neural Interactive Collaborative Filtering 4 Jul 2020 · 1 repository · arXiv:2007.02095