Methods › Reinforcement Learning › Off-Policy TD Control › Q-Learning › Papers, page 10
Q-Learning
Papers archive 2025-07-28
archive papers tagged: 1,734 · with a code link: 464 · where Syntology ran a sample: 126 (105 with a run with no instrument failure, 21 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (126 of 1,734 tagged: 105 with a run with no instrument failure, 21 where every run was a failure of Syntology's instrument)
Page 10 of 18: papers 901 to 1,000 of 1,734, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Density Estimation for Conservative Q-Learning 29 Sep 2021 · 0 repositories
-
Disentangling Generalization in Reinforcement Learning 29 Sep 2021 · 0 repositories
-
Explanation-Aware Experience Replay in Rule-Dense Environments 29 Sep 2021 · 1 repository · arXiv:2109.14711
-
HyperDQN: A Randomized Exploration Method for Deep Reinforcement Learning 29 Sep 2021 · 1 repository
-
Learning Explicit Credit Assignment for Multi-agent Joint Q-learning 29 Sep 2021 · 0 repositories
-
Offline Reinforcement Learning with In-sample Q-Learning 29 Sep 2021 · 1 repository
-
On the Estimation Bias in Double Q-Learning 29 Sep 2021 · 1 repository · arXiv:2109.14419Syntology official (archive's flag): 4 ran · 4 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples)
-
Online Robust Reinforcement Learning with Model Uncertainty 29 Sep 2021 · 0 repositories · arXiv:2109.14523
-
Q-learning for real time control of heterogeneous microagent collectives 29 Sep 2021 · 0 repositories
-
Q-Learning Scheduler for Multi-Task Learning through the use of Histogram of Task Uncertainty 29 Sep 2021 · 0 repositories
-
Robust and Data-efficient Q-learning by Composite Value-estimation 29 Sep 2021 · 0 repositories
-
δ²-exploration for Reinforcement Learning 29 Sep 2021 · 0 repositories
-
Text Generation with Efficient (Soft) Q-Learning 29 Sep 2021 · 0 repositories
-
The guide and the explorer: smart agents for resource-limited iterated batch reinforcement learning 29 Sep 2021 · 0 repositories
-
Towards Unknown-aware Deep Q-Learning 29 Sep 2021 · 0 repositories
-
Unifying Top-down and Bottom-up for Recurrent Visual Attention 29 Sep 2021 · 0 repositories
-
Deep Reinforcement Learning with Adjustments 28 Sep 2021 · 0 repositories · arXiv:2109.13463
-
Smart Home Energy Management: Sequence-to-Sequence Load Forecasting and Q-Learning 25 Sep 2021 · 0 repositories · arXiv:2109.12440
-
Parameter-free Reduction of the Estimation Bias in Deep Reinforcement Learning for Deterministic Policy Gradients 24 Sep 2021 · 1 repository · arXiv:2109.11788
-
Fetal oxygen delivery and consumption and blood gases in relation to gestational age 23 Sep 2021 · 0 repositories · arXiv:2109.11616
-
Estimation Error Correction in Deep Reinforcement Learning for Deterministic Actor-Critic Methods 22 Sep 2021 · 1 repository · arXiv:2109.10736
-
MEPG: A Minimalist Ensemble Policy Gradient Framework for Deep Reinforcement Learning 22 Sep 2021 · 0 repositories · arXiv:2109.10552
-
Off-line approximate dynamic programming for the vehicle routing problem with a highly variable customer basis and stochastic demands 21 Sep 2021 · 0 repositories · arXiv:2109.10200
-
Greedy UnMixing for Q-Learning in Multi-Agent Reinforcement Learning 19 Sep 2021 · 0 repositories · arXiv:2109.09034
-
Learning from Peers: Deep Transfer Reinforcement Learning for Joint Radio and Cache Resource Allocation in 5G RAN Slicing 16 Sep 2021 · 0 repositories · arXiv:2109.07999
-
Reinforcement Learning on Encrypted Data 16 Sep 2021 · 0 repositories · arXiv:2109.08236
-
Convergence of a Human-in-the-Loop Policy-Gradient Algorithm With Eligibility Trace Under Reward, Policy, and Advantage Feedback 15 Sep 2021 · 0 repositories · arXiv:2109.07054
-
Optimal Cycling of a Heterogenous Battery Bank via Reinforcement Learning 15 Sep 2021 · 0 repositories · arXiv:2109.07137
-
Deep hierarchical reinforcement agents for automated penetration testing 14 Sep 2021 · 0 repositories · arXiv:2109.06449
-
Vision Transformer for Learning Driving Policies in Complex Multi-Agent Environments 14 Sep 2021 · 0 repositories · arXiv:2109.06514
-
Bootstrapped Meta-Learning 9 Sep 2021 · 1 repository · arXiv:2109.04504
-
Learning cortical representations through perturbed and adversarial dreaming 9 Sep 2021 · 1 repository · arXiv:2109.04261Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
User Tampering in Reinforcement Learning Recommender Systems 9 Sep 2021 · 0 repositories · arXiv:2109.04083
-
Deep SIMBAD: Active Landmark-based Self-localization Using Ranking -based Scene Descriptor 6 Sep 2021 · 0 repositories · arXiv:2109.02786
-
Learning-Based Strategy Design for Robot-Assisted Reminiscence Therapy Based on a Developed Model for People with Dementia 6 Sep 2021 · 0 repositories · arXiv:2109.02194
-
Temporal Shift Reinforcement Learning 5 Sep 2021 · 1 repository · arXiv:2109.02145
-
Event-Based Communication in Distributed Q-Learning 3 Sep 2021 · 0 repositories · arXiv:2109.01417
-
Deep Reinforcement Learning at the Edge of the Statistical Precipice 30 Aug 2021 · 3 repositories · arXiv:2108.13264Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 4 honoured, 1 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
Deep Reinforcement Learning for Dynamic Band Switch in Cellular-Connected UAV 26 Aug 2021 · 0 repositories · arXiv:2108.12054
-
DQLEL: Deep Q-Learning for Energy-Optimized LoS/NLoS UWB Node Selection 24 Aug 2021 · 0 repositories · arXiv:2108.13157
-
An Independent Study of Reinforcement Learning and Autonomous Driving 20 Aug 2021 · 0 repositories · arXiv:2110.07729
-
A Microscopic Pandemic Simulator for Pandemic Prediction Using Scalable Million-Agent Reinforcement Learning 14 Aug 2021 · 0 repositories · arXiv:2108.06589
-
DQN Control Solution for KDD Cup 2021 City Brain Challenge 14 Aug 2021 · 1 repository · arXiv:2108.06491
-
DQ-GAT: Towards Safe and Efficient Autonomous Driving with Deep Q-Learning and Graph Attention Networks 11 Aug 2021 · 0 repositories · arXiv:2108.05030
-
Two is a crowd: tracking relations in videos 11 Aug 2021 · 0 repositories · arXiv:2108.05331
-
Modified Double DQN: addressing stability 9 Aug 2021 · 0 repositories · arXiv:2108.04115
-
An Elementary Proof that Q-learning Converges Almost Surely 5 Aug 2021 · 0 repositories · arXiv:2108.02827
-
Offline Decentralized Multi-Agent Reinforcement Learning 4 Aug 2021 · 0 repositories · arXiv:2108.01832
-
SABER: Data-Driven Motion Planner for Autonomously Navigating Heterogeneous Robots 3 Aug 2021 · 1 repository · arXiv:2108.01262
-
An Efficient Image-to-Image Translation HourGlass-based Architecture for Object Pushing Policy Learning 2 Aug 2021 · 1 repository · arXiv:2108.01034
-
A DQN-based Approach to Finding Precise Evidences for Fact Verification 1 Aug 2021 · 1 repository
-
REM: Efficient Semi-Automated Real-Time Moderation of Online Forums 1 Aug 2021 · 0 repositories
-
A Distributed Intelligence Architecture for B5G Network Automation 28 Jul 2021 · 0 repositories · arXiv:2107.13268
-
An Improved Algorithm of Robot Path Planning in Complex Environment Based on Double DQN 23 Jul 2021 · 0 repositories · arXiv:2107.11245
-
Constraints Penalized Q-learning for Safe Offline Reinforcement Learning 19 Jul 2021 · 0 repositories · arXiv:2107.09003
-
A Reinforcement Learning Environment for Mathematical Reasoning via Program Synthesis 15 Jul 2021 · 1 repository · arXiv:2107.07373
-
Deep Reinforcement Learning based Dynamic Optimization of Bus Timetable 15 Jul 2021 · 0 repositories · arXiv:2107.07066
-
Minimizing Safety Interference for Safe and Comfortable Automated Driving with Distributional Reinforcement Learning 15 Jul 2021 · 0 repositories · arXiv:2107.07316
-
A Penalized Shared-parameter Algorithm for Estimating Optimal Dynamic Treatment Regimens 13 Jul 2021 · 0 repositories · arXiv:2107.07875
-
Transfer Learning in Multi-Agent Reinforcement Learning with Double Q-Networks for Distributed Resource Sharing in V2X Communication 13 Jul 2021 · 0 repositories · arXiv:2107.06195
-
Backprop-Free Reinforcement Learning with Active Neural Generative Coding 10 Jul 2021 · 1 repository · arXiv:2107.07046Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples)
-
Reinforced Hybrid Genetic Algorithm for the Traveling Salesman Problem 9 Jul 2021 · 0 repositories · arXiv:2107.06870
-
Computational Benefits of Intermediate Rewards for Goal-Reaching Policy Learning 8 Jul 2021 · 1 repository · arXiv:2107.03961Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 1 pointer-only (licence)
-
Ensemble and Auxiliary Tasks for Data-Efficient Deep Reinforcement Learning 5 Jul 2021 · 1 repository · arXiv:2107.01904Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Improve Agents without Retraining: Parallel Tree Search with Off-Policy Correction 4 Jul 2021 · 1 repository · arXiv:2107.01715
-
A Novel Deep Reinforcement Learning Based Stock Direction Prediction using Knowledge Graph and Community Aware Sentiments 2 Jul 2021 · 0 repositories · arXiv:2107.00931
-
AoI Minimization in Energy Harvesting and Spectrum Sharing Enabled 6G Networks 1 Jul 2021 · 0 repositories · arXiv:2107.00340
-
Distilling Reinforcement Learning Tricks for Video Games 1 Jul 2021 · 1 repository · arXiv:2107.00703
-
Gap-Dependent Bounds for Two-Player Markov Games 1 Jul 2021 · 0 repositories · arXiv:2107.00685
-
Markov Decision Process modeled with Bandits for Sequential Decision Making in Linear-flow 1 Jul 2021 · 0 repositories · arXiv:2107.00204
-
Convergent and Efficient Deep Q Network Algorithm 29 Jun 2021 · 1 repository · arXiv:2106.15419
-
DRILL-- Deep Reinforcement Learning for Refinement Operators in 𝒜ℒ𝒞 29 Jun 2021 · 0 repositories · arXiv:2106.15373
-
Expert Q-learning: Deep Reinforcement Learning with Coarse State Values from Offline Expert Examples 28 Jun 2021 · 0 repositories · arXiv:2106.14642
-
Concentration of Contractive Stochastic Approximation and Reinforcement Learning 27 Jun 2021 · 0 repositories · arXiv:2106.14308
-
Reinforcement Learning for Mean Field Games, with Applications to Economics 25 Jun 2021 · 0 repositories · arXiv:2106.13755
-
Exploration-Exploitation in Multi-Agent Competition: Convergence with Bounded Rationality 24 Jun 2021 · 0 repositories · arXiv:2106.12928
-
Coarse-to-Fine Q-attention: Efficient Learning for Visual Robotic Manipulation via Discretisation 23 Jun 2021 · 1 repository · arXiv:2106.12534
-
MMD-MIX: Value Function Factorisation with Maximum Mean Discrepancy for Cooperative Multi-Agent Reinforcement Learning 22 Jun 2021 · 0 repositories · arXiv:2106.11652
-
Q-Learning Lagrange Policies for Multi-Action Restless Bandits 22 Jun 2021 · 1 repository · arXiv:2106.12024Syntology official (archive's flag): 14 ran · 14 ran (of which 0 constructed an object rather than computing a result; 14 with no instrument failure: 0 honoured, 0 violated, 14 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 17 harvested samples)
-
Reinforcement Learning for Physical Layer Communications 22 Jun 2021 · 1 repository · arXiv:2106.11595
-
Analytically Tractable Bayesian Deep Q-Learning 21 Jun 2021 · 0 repositories · arXiv:2106.11086
-
Distributed Heuristic Multi-Agent Path Finding with Communication 21 Jun 2021 · 1 repository · arXiv:2106.11365
-
Reinforcement Learning for Resource Allocation in Steerable Laser-based Optical Wireless Systems 21 Jun 2021 · 0 repositories · arXiv:2106.11368
-
A Deep Reinforcement Learning Approach towards Pendulum Swing-up Problem based on TF-Agents 17 Jun 2021 · 0 repositories · arXiv:2106.09556
-
Unbiased Methods for Multi-Goal Reinforcement Learning 16 Jun 2021 · 0 repositories · arXiv:2106.08863
-
Vision-Language Navigation with Random Environmental Mixup 15 Jun 2021 · 1 repository · arXiv:2106.07876Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 4 harvested samples)
-
Efficient (Soft) Q-Learning for Text Generation with Limited Good Data 14 Jun 2021 · 1 repository · arXiv:2106.07704Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
DouZero: Mastering DouDizhu with Self-Play Deep Reinforcement Learning 11 Jun 2021 · 1 repository · arXiv:2106.06135Syntology official (archive's flag): 2 ran · 3 ran (of which 1 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 6 harvested samples) · 1 pointer-only (licence)
-
GDI: Rethinking What Makes Reinforcement Learning Different From Supervised Learning 11 Jun 2021 · 0 repositories · arXiv:2106.06232
-
Reinforced Few-Shot Acquisition Function Learning for Bayesian Optimization 8 Jun 2021 · 0 repositories · arXiv:2106.04335
-
Believe What You See: Implicit Constraint Approach for Offline Multi-Agent Reinforcement Learning 7 Jun 2021 · 1 repository · arXiv:2106.03400
-
Decentralized Q-Learning in Zero-sum Markov Games 4 Jun 2021 · 0 repositories · arXiv:2106.02748
-
Design and Comparison of Reward Functions in Reinforcement Learning for Energy Management of Sensor Nodes 2 Jun 2021 · 0 repositories · arXiv:2106.01114
-
Smooth Q-learning: Accelerate Convergence of Q-learning Using Similarity 2 Jun 2021 · 0 repositories · arXiv:2106.01134
-
Energy-aware optimization of UAV base stations placement via decentralized multi-agent Q-learning 1 Jun 2021 · 0 repositories · arXiv:2106.00845
-
SHAQ: Incorporating Shapley Value Theory into Multi-Agent Q-Learning 31 May 2021 · 1 repository · arXiv:2105.15013Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Sample-Efficient Reinforcement Learning for Linearly-Parameterized MDPs with a Generative Model 28 May 2021 · 0 repositories · arXiv:2105.14016
-
Learning to Optimize Industry-Scale Dynamic Pickup and Delivery Problems 27 May 2021 · 0 repositories · arXiv:2105.12899
-
Reputation Bootstrapping for Composite Services using CP-nets 27 May 2021 · 0 repositories · arXiv:2105.15135
-
Verification of Dissipativity and Evaluation of Storage Function in Economic Nonlinear MPC using Q-Learning 24 May 2021 · 0 repositories · arXiv:2105.11313