Methods › Reinforcement Learning › Off-Policy TD Control › Q-Learning › Papers, page 14
Q-Learning
Papers archive 2025-07-28
archive papers tagged: 1,734 · with a code link: 464 · where Syntology ran a sample: 126 (105 with a run with no instrument failure, 21 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (126 of 1,734 tagged: 105 with a run with no instrument failure, 21 where every run was a failure of Syntology's instrument)
Page 14 of 18: papers 1,301 to 1,400 of 1,734, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Periodic Q-Learning 23 Feb 2020 · 0 repositories · arXiv:2002.09795
-
Anypath Routing Protocol Design via Q-Learning for Underwater Sensor Networks 22 Feb 2020 · 0 repositories · arXiv:2002.09623
-
Disentangling Controllable Object through Video Prediction Improves Visual Reinforcement Learning 21 Feb 2020 · 0 repositories · arXiv:2002.09136
-
Langevin DQN 17 Feb 2020 · 2 repositories · arXiv:2002.07282
-
A Multimodal Dialogue System for Conversational Image Editing 16 Feb 2020 · 0 repositories · arXiv:2002.06484
-
Maxmin Q-learning: Controlling the Estimation Bias of Q-learning 16 Feb 2020 · 1 repository · arXiv:2002.06487
-
Reinforced active learning for image segmentation 16 Feb 2020 · 1 repository · arXiv:2002.06583
-
Fast Reinforcement Learning for Anti-jamming Communications 13 Feb 2020 · 0 repositories · arXiv:2002.05364
-
Listwise Learning to Rank with Deep Q-Networks 13 Feb 2020 · 0 repositories · arXiv:2002.07651
-
Regret Bounds for Discounted MDPs 12 Feb 2020 · 0 repositories · arXiv:2002.05138
-
Mean-Field Controls with Q-learning for Cooperative MARL: Convergence and Complexity Analysis 10 Feb 2020 · 0 repositories · arXiv:2002.04131
-
Learning State Abstractions for Transfer in Continuous Control 8 Feb 2020 · 2 repositories · arXiv:2002.05518
-
Safe Wasserstein Constrained Deep Q-Learning 7 Feb 2020 · 0 repositories · arXiv:2002.03016
-
Deep Radial-Basis Value Functions for Continuous Control 5 Feb 2020 · 0 repositories · arXiv:2002.01883
-
Interpretable End-to-end Urban Autonomous Driving with Latent Deep Reinforcement Learning 23 Jan 2020 · 4 repositories · arXiv:2001.08726Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 6 harvested samples)
-
Q-Learning in enormous action spaces via amortized approximate maximization 22 Jan 2020 · 0 repositories · arXiv:2001.08116
-
Discriminator Soft Actor Critic without Extrinsic Rewards 19 Jan 2020 · 1 repository · arXiv:2001.06808
-
Model-based Multi-Agent Reinforcement Learning with Cooperative Prioritized Sweeping 15 Jan 2020 · 0 repositories · arXiv:2001.07527
-
Deep Interactive Reinforcement Learning for Path Following of Autonomous Underwater Vehicle 10 Jan 2020 · 0 repositories · arXiv:2001.03359
-
A Probabilistic Simulator of Spatial Demand for Product Allocation 9 Jan 2020 · 0 repositories · arXiv:2001.03210
-
EEG-based Drowsiness Estimation for Driving Safety using Deep Q-Learning 8 Jan 2020 · 0 repositories · arXiv:2001.02399
-
Experimental Analysis of Reinforcement Learning Techniques for Spectrum Sharing Radar 6 Jan 2020 · 0 repositories · arXiv:2001.01799
-
Deep Randomized Least Squares Value Iteration 1 Jan 2020 · 0 repositories
-
SVQN: Sequential Variational Soft Q-Learning Networks 1 Jan 2020 · 0 repositories
-
Way Off-Policy Batch Deep Reinforcement Learning of Human Preferences in Dialog 1 Jan 2020 · 0 repositories
-
Information Theoretic Model Predictive Q-Learning 31 Dec 2019 · 0 repositories · arXiv:2001.02153
-
SLM Lab: A Comprehensive Benchmark and Modular Software Framework for Reproducible Deep Reinforcement Learning 28 Dec 2019 · 1 repository · arXiv:1912.12482Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 13 harvested samples)
-
Hamilton-Jacobi-Bellman Equations for Q-Learning in Continuous Time 23 Dec 2019 · 0 repositories · arXiv:1912.10697
-
Learning an Interpretable Traffic Signal Control Policy 23 Dec 2019 · 1 repository · arXiv:1912.11023
-
Exploiting the potential of deep reinforcement learning for classification tasks in high-dimensional and unstructured data 20 Dec 2019 · 0 repositories · arXiv:1912.09595
-
Soft Q Network 20 Dec 2019 · 0 repositories · arXiv:1912.10891
-
Sepsis World Model: A MIMIC-based OpenAI Gym "World Model" Simulator for Sepsis Treatment 15 Dec 2019 · 0 repositories · arXiv:1912.07127
-
High dimensional precision medicine from patient-derived xenografts 13 Dec 2019 · 0 repositories · arXiv:1912.06667
-
Provably Efficient Reinforcement Learning with Aggregated States 13 Dec 2019 · 0 repositories · arXiv:1912.06366
-
A Finite-Time Analysis of Q-Learning with Neural Network Function Approximation 10 Dec 2019 · 0 repositories · arXiv:1912.04511
-
Learning Sparse Representations Incrementally in Deep Reinforcement Learning 9 Dec 2019 · 0 repositories · arXiv:1912.04002
-
Value-of-Information based Arbitration between Model-based and Model-free Control 8 Dec 2019 · 0 repositories · arXiv:1912.05453
-
Reinforcement Learning with Non-Markovian Rewards 5 Dec 2019 · 0 repositories · arXiv:1912.02552
-
Combining Q-Learning and Search with Amortized Value Estimates 5 Dec 2019 · 0 repositories · arXiv:1912.02807
-
A Unified Switching System Perspective and O.D.E. Analysis of Q-Learning Algorithms 4 Dec 2019 · 0 repositories · arXiv:1912.02270
-
Learning to Dynamically Coordinate Multi-Robot Teams in Graph Attention Networks 4 Dec 2019 · 0 repositories · arXiv:1912.02059
-
Neighborhood Cognition Consistent Multi-Agent Reinforcement Learning 3 Dec 2019 · 0 repositories · arXiv:1912.01160
-
Modelling the Dynamics of Multiagent Q-Learning in Repeated Symmetric Games: a Mean Field Theoretic Approach 1 Dec 2019 · 0 repositories
-
Propagating Uncertainty in Reinforcement Learning via Wasserstein Barycenters 1 Dec 2019 · 1 repository
-
Provably Efficient Q-learning with Function Approximation via Distribution Shift Error Checking Oracle 1 Dec 2019 · 0 repositories
-
Reconciling λ-Returns with Experience Replay 1 Dec 2019 · 1 repository
-
Quadratic Q-network for Learning Continuous Control for Autonomous Vehicles 29 Nov 2019 · 0 repositories · arXiv:1912.00074
-
Control-Tutored Reinforcement Learning: an application to the Herding Problem 26 Nov 2019 · 0 repositories · arXiv:1911.11444
-
Join Query Optimization with Deep Reinforcement Learning Algorithms 26 Nov 2019 · 1 repository · arXiv:1911.11689
-
Adaptive Modulation and Coding based on Reinforcement Learning for 5G Networks 25 Nov 2019 · 0 repositories · arXiv:1912.04030
-
Mitigate Bias in Face Recognition using Skewness-Aware Reinforcement Learning 25 Nov 2019 · 0 repositories · arXiv:1911.10692
-
Which Channel to Ask My Question? Personalized Customer Service RequestStream Routing using DeepReinforcement Learning 24 Nov 2019 · 0 repositories · arXiv:1911.10521
-
Efficient Drone Mobility Support Using Reinforcement Learning 21 Nov 2019 · 0 repositories · arXiv:1911.09715
-
Quantum Observables for continuous control of the Quantum Approximate Optimization Algorithm via Reinforcement Learning 21 Nov 2019 · 0 repositories · arXiv:1911.09682
-
Placement Optimization of Aerial Base Stations with Deep Reinforcement Learning 19 Nov 2019 · 0 repositories · arXiv:1911.08111
-
Asymptotics of Reinforcement Learning with Neural Networks 13 Nov 2019 · 0 repositories · arXiv:1911.07304
-
Minimalistic Attacks: How Little it Takes to Fool a Deep Reinforcement Learning Policy 10 Nov 2019 · 0 repositories · arXiv:1911.03849
-
Two-stage WECC Composite Load Modeling: A Double Deep Q-Learning Networks Approach 8 Nov 2019 · 0 repositories · arXiv:1911.04894
-
An End-to-End Deep RL Framework for Task Arrangement in Crowdsourcing Platforms 4 Nov 2019 · 0 repositories · arXiv:1911.01030
-
Challenging On Car Racing Problem from OpenAI gym 2 Nov 2019 · 0 repositories · arXiv:1911.04868
-
On Solving the 2-Dimensional Greedy Shooter Problem for UAVs 2 Nov 2019 · 1 repository · arXiv:1911.01419
-
Generalized Speedy Q-learning 1 Nov 2019 · 1 repository · arXiv:1911.00397
-
Biomimetic Ultra-Broadband Perfect Absorbers Optimised with Reinforcement Learning 28 Oct 2019 · 0 repositories · arXiv:1910.12465
-
Model-Free Mean-Field Reinforcement Learning: Mean-Field MDP and Mean-Field Q-Learning 28 Oct 2019 · 0 repositories · arXiv:1910.12802
-
BAIL: Best-Action Imitation Learning for Batch Deep Reinforcement Learning 27 Oct 2019 · 1 repository · arXiv:1910.12179Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Task-Oriented Language Grounding for Language Input with Multiple Sub-Goals of Non-Linear Order 27 Oct 2019 · 1 repository · arXiv:1910.12354
-
D-Point Trigonometric Path Planning based on Q-Learning in Uncertain Environments 26 Oct 2019 · 0 repositories · arXiv:1910.12020
-
ZPD Teaching Strategies for Deep Reinforcement Learning from Demonstrations 26 Oct 2019 · 2 repositories · arXiv:1910.12154Syntology 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples)
-
Deep Q-Learning for Same-Day Delivery with Vehicles and Drones 25 Oct 2019 · 0 repositories · arXiv:1910.11901
-
Momentum in Reinforcement Learning 21 Oct 2019 · 0 repositories · arXiv:1910.09322
-
Resource Allocation in Mobility-Aware Federated Learning Networks: A Deep Reinforcement Learning Approach 21 Oct 2019 · 0 repositories · arXiv:1910.09172
-
Policy Learning for Malaria Control 20 Oct 2019 · 2 repositories · arXiv:1910.08926
-
Reverse Experience Replay 19 Oct 2019 · 0 repositories · arXiv:1910.08780
-
Automatic Data Augmentation by Learning the Deterministic Policy 18 Oct 2019 · 1 repository · arXiv:1910.08343
-
On the Reduction of Variance and Overestimation of Deep Q-Learning 14 Oct 2019 · 0 repositories · arXiv:1910.05983
-
Zap Q-Learning With Nonlinear Function Approximation 11 Oct 2019 · 0 repositories · arXiv:1910.05405
-
Integrating Behavior Cloning and Reinforcement Learning for Improved Performance in Dense and Sparse Reward Environments 9 Oct 2019 · 0 repositories · arXiv:1910.04281
-
Learning Visual Affordances with Target-Orientated Deep Q-Network to Grasp Objects by Harnessing Environmental Fixtures 9 Oct 2019 · 0 repositories · arXiv:1910.03781
-
Toward Synergic Learning for Autonomous Manipulation of Deformable Tissues via Surgical Robots: An Approximate Q-Learning Approach 8 Oct 2019 · 0 repositories · arXiv:1910.03398
-
Multi-step Greedy Reinforcement Learning Algorithms 7 Oct 2019 · 0 repositories · arXiv:1910.02919
-
Reinforcement Learning with Structured Hierarchical Grammar Representations of Actions 7 Oct 2019 · 0 repositories · arXiv:1910.02876
-
Deep Q-Network for Angry Birds 4 Oct 2019 · 1 repository · arXiv:1910.01806
-
I'm sorry Dave, I'm afraid I can't do that, Deep Q-learning from forbidden action 4 Oct 2019 · 0 repositories · arXiv:1910.02078
-
Benchmarking Batch Deep Reinforcement Learning Algorithms 3 Oct 2019 · 5 repositories · arXiv:1910.01708
-
AI Assisted Annotator using Reinforcement Learning 2 Oct 2019 · 0 repositories · arXiv:1910.02052
-
Quantile QT-Opt for Risk-Aware Vision-Based Robotic Grasping 1 Oct 2019 · 0 repositories · arXiv:1910.02787
-
Meta-Q-Learning 30 Sep 2019 · 2 repositories · arXiv:1910.00125Syntology community repositories only · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Composite Q-learning: Multi-scale Q-function Decomposition and Separable Optimization 30 Sep 2019 · 0 repositories · arXiv:1909.13518
-
Visual Exploration and Energy-aware Path Planning via Reinforcement Learning 26 Sep 2019 · 1 repository · arXiv:1909.12217
-
CAQL: Continuous Action Q-Learning 26 Sep 2019 · 0 repositories · arXiv:1909.12397
-
Active inference: demystified and compared 24 Sep 2019 · 1 repository · arXiv:1909.10863
-
Deep Reinforcement Learning with Modulated Hebbian plus Q Network Architecture 21 Sep 2019 · 1 repository · arXiv:1909.09902
-
On the Convergence of Approximate and Regularized Policy Iteration Schemes 20 Sep 2019 · 0 repositories · arXiv:1909.09621
-
ModelicaGym: Applying Reinforcement Learning to Modelica Models 18 Sep 2019 · 1 repository · arXiv:1909.08604
-
Split Deep Q-Learning for Robust Object Singulation 17 Sep 2019 · 0 repositories · arXiv:1909.08105
-
Joint Inference of Reward Machines and Policies for Reinforcement Learning 12 Sep 2019 · 0 repositories · arXiv:1909.05912
-
Mutual-Information Regularization in Markov Decision Processes and Actor-Critic Learning 11 Sep 2019 · 0 repositories · arXiv:1909.05950
-
Q-learning Assisted Energy-Aware Traffic Offloading and Cell Switching in Heterogeneous Networks 11 Sep 2019 · 0 repositories
-
A Multistep Lyapunov Approach for Finite-Time Analysis of Biased Stochastic Approximation 10 Sep 2019 · 0 repositories · arXiv:1909.04299
-
Fixed-Horizon Temporal Difference Methods for Stable Reinforcement Learning 9 Sep 2019 · 0 repositories · arXiv:1909.03906