Methods › Reinforcement Learning › Off-Policy TD Control › Q-Learning › Papers, page 11
Q-Learning
Papers archive 2025-07-28
archive papers tagged: 1,734 · with a code link: 464 · where Syntology ran a sample: 126 (105 with a run with no instrument failure, 21 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (126 of 1,734 tagged: 105 with a run with no instrument failure, 21 where every run was a failure of Syntology's instrument)
Page 11 of 18: papers 1,001 to 1,100 of 1,734, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Deep Reinforcement Learning for Optimal Stopping with Application in Financial Engineering 19 May 2021 · 1 repository · arXiv:2105.08877
-
Improved Exploring Starts by Kernel Density Estimation-Based State-Space Coverage Acceleration in Reinforcement Learning 19 May 2021 · 1 repository · arXiv:2105.08990
-
Efficient Off-Policy Q-Learning for Data-Based Discrete-Time LQR Problems 17 May 2021 · 0 repositories · arXiv:2105.07761
-
Learn to Intervene: An Adaptive Learning Policy for Restless Bandits in Application to Preventive Healthcare 17 May 2021 · 0 repositories · arXiv:2105.07965
-
Uncertainty Weighted Actor-Critic for Offline Reinforcement Learning 17 May 2021 · 2 repositories · arXiv:2105.08140
-
Interpretable performance analysis towards offline reinforcement learning: A dataset perspective 12 May 2021 · 0 repositories · arXiv:2105.05473
-
Fast constraint satisfaction problem and learning-based algorithm for solving Minesweeper 10 May 2021 · 0 repositories · arXiv:2105.04120
-
Reinforcement Learning with Expert Trajectory For Quantitative Trading 9 May 2021 · 0 repositories · arXiv:2105.03844
-
Time-Aware Q-Networks: Resolving Temporal Irregularity for Deep Reinforcement Learning 6 May 2021 · 0 repositories · arXiv:2105.02580
-
Automated scoring of pre-REM sleep in mice with deep learning 5 May 2021 · 1 repository · arXiv:2105.01933
-
Survey on Multi-Agent Q-Learning frameworks for resource management in wireless sensor network 5 May 2021 · 0 repositories · arXiv:2105.02371
-
HASCO: Towards Agile HArdware and Software CO-design for Tensor Computation 4 May 2021 · 1 repository · arXiv:2105.01585
-
Action Candidate Based Clipped Double Q-learning for Discrete and Continuous Action Tasks 3 May 2021 · 1 repository · arXiv:2105.00704
-
Robotic Surgery With Lean Reinforcement Learning 3 May 2021 · 1 repository · arXiv:2105.01006
-
CARL-DTN: Context Adaptive Reinforcement Learning based Routing Algorithm in Delay Tolerant Network 2 May 2021 · 0 repositories · arXiv:2105.00544
-
Adapting to Reward Progressivity via Spectral Reinforcement Learning 29 Apr 2021 · 1 repository · arXiv:2104.14138Syntology official (archive's flag): 2 ran · 2 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; every one of the 2 samples that ran constructed an object rather than computing a result (of 2 harvested samples) · 2 pointer-only (licence)
-
Emotional Contagion-Aware Deep Reinforcement Learning for Antagonistic Crowd Simulation 29 Apr 2021 · 0 repositories · arXiv:2105.00854
-
RP-DQN: An application of Q-Learning to Vehicle Routing Problems 25 Apr 2021 · 0 repositories · arXiv:2104.12226
-
Independent Reinforcement Learning for Weakly Cooperative Multiagent Traffic Control Problem 22 Apr 2021 · 1 repository · arXiv:2104.10917
-
Model-aided Deep Reinforcement Learning for Sample-efficient UAV Trajectory Design in IoT Networks 21 Apr 2021 · 0 repositories · arXiv:2104.10403
-
Reinforcement Learning for Traffic Signal Control: Comparison with Commercial Systems 21 Apr 2021 · 0 repositories · arXiv:2104.10455
-
A Simulated Experiment to Explore Robotic Dialogue Strategies for People with Dementia 18 Apr 2021 · 0 repositories · arXiv:2104.08940
-
Reinforcement learning based process optimization and strategy development in conventional tunneling 17 Apr 2021 · 1 repository
-
Actionable Models: Unsupervised Offline Reinforcement Learning of Robotic Skills 15 Apr 2021 · 0 repositories · arXiv:2104.07749
-
A coevolutionary approach to deep multi-agent reinforcement learning 12 Apr 2021 · 1 repository · arXiv:2104.05610
-
Prospect-theoretic Q-learning 12 Apr 2021 · 0 repositories · arXiv:2104.05311
-
Autoequivariant Network Search via Group Decomposition 10 Apr 2021 · 1 repository · arXiv:2104.04848
-
Optimal Market Making by Reinforcement Learning 8 Apr 2021 · 1 repository · arXiv:2104.04036
-
Distributed Deep Reinforcement Learning for Collaborative Spectrum Sharing 6 Apr 2021 · 0 repositories · arXiv:2104.02059
-
SOLO: Search Online, Learn Offline for Combinatorial Optimization Problems 4 Apr 2021 · 0 repositories · arXiv:2104.01646
-
Convergence of Finite Memory Q-Learning for POMDPs and Near Optimality of Learned Policies under Filter Stability 22 Mar 2021 · 0 repositories · arXiv:2103.12158
-
Reinforcement Learning based on Scenario-tree MPC for ASVs 22 Mar 2021 · 0 repositories · arXiv:2103.11949
-
Variational quantum compiling with double Q-learning 22 Mar 2021 · 0 repositories · arXiv:2103.11611
-
Full Gradient DQN Reinforcement Learning: A Provably Convergent Scheme 10 Mar 2021 · 0 repositories · arXiv:2103.05981
-
S4RL: Surprisingly Simple Self-Supervision for Offline Reinforcement Learning 10 Mar 2021 · 0 repositories · arXiv:2103.06326
-
Increasing Energy Efficiency of Massive-MIMO Network via Base Stations Switching using Reinforcement Learning and Radio Environment Maps 8 Mar 2021 · 0 repositories · arXiv:2103.11891
-
Correlated Deep Q-learning based Microgrid Energy Management 6 Mar 2021 · 0 repositories · arXiv:2103.04152
-
Decentralized Microgrid Energy Management: A Multi-agent Correlated Q-learning Approach 6 Mar 2021 · 0 repositories · arXiv:2103.04154
-
Super-resolving Compressed Images via Parallel and Series Integration of Artifact Reduction and Resolution Enhancement 2 Mar 2021 · 1 repository · arXiv:2103.01698
-
UCB Momentum Q-learning: Correcting the bias without forgetting 1 Mar 2021 · 1 repository · arXiv:2103.01312
-
Ensemble Bootstrapping for Q-Learning 28 Feb 2021 · 0 repositories · arXiv:2103.00445
-
Multi-Agent Path Planning based on MPC and DDPG 26 Feb 2021 · 0 repositories · arXiv:2102.13283
-
Balancing Rational and Other-Regarding Preferences in Cooperative-Competitive Environments 24 Feb 2021 · 1 repository · arXiv:2102.12307
-
Sequential Learning-based IaaS Composition 24 Feb 2021 · 0 repositories · arXiv:2102.12598
-
Greedy-Step Off-Policy Reinforcement Learning 23 Feb 2021 · 0 repositories · arXiv:2102.11717
-
Stratified Experience Replay: Correcting Multiplicity Bias in Off-Policy Reinforcement Learning 22 Feb 2021 · 0 repositories · arXiv:2102.11319
-
Training a Resilient Q-Network against Observational Interference 18 Feb 2021 · 1 repository · arXiv:2102.09677
-
Adaptive Rational Activations to Boost Deep Reinforcement Learning 18 Feb 2021 · 4 repositories · arXiv:2102.09407Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 6 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples) · 3 pointer-only (licence)
-
A Discrete-Time Switching System Analysis of Q-learning 17 Feb 2021 · 0 repositories · arXiv:2102.08583
-
DFAC Framework: Factorizing the Value Function via Quantile Mixture for Multi-Agent Distributional Q-Learning 16 Feb 2021 · 1 repository · arXiv:2102.07936Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample)
-
Cooperation and Reputation Dynamics with Reinforcement Learning 15 Feb 2021 · 0 repositories · arXiv:2102.07523
-
Reversible Action Design for Combinatorial Optimization with Reinforcement Learning 14 Feb 2021 · 0 repositories · arXiv:2102.07210
-
Is Q-Learning Minimax Optimal? A Tight Sample Complexity Analysis 12 Feb 2021 · 0 repositories · arXiv:2102.06548
-
Deep Reinforcement Agent for Scheduling in HPC 11 Feb 2021 · 1 repository · arXiv:2102.06243
-
Benchmarking Deep Graph Generative Models for Optimizing New Drug Molecules for COVID-19 9 Feb 2021 · 1 repository · arXiv:2102.04977
-
A review of motion planning algorithms for intelligent robotics 4 Feb 2021 · 0 repositories · arXiv:2102.02376
-
A step toward a reinforcement learning de novo genome assembler 2 Feb 2021 · 0 repositories · arXiv:2102.02649
-
Acting in Delayed Environments with Non-Stationary Markov Policies 28 Jan 2021 · 2 repositories · arXiv:2101.11992Syntology official (archive's flag): 3 ran · 3 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 4 unverified (of 7 harvested samples)
-
Reinforcement Learning based Per-antenna Discrete Power Control for Massive MIMO Systems 28 Jan 2021 · 0 repositories · arXiv:2101.12154
-
Reinforcement Learning Assisted Beamforming for Inter-cell Interference Mitigation in 5G Massive MIMO Networks 27 Jan 2021 · 0 repositories · arXiv:2103.11782
-
Channel Estimation via Successive Denoising in MIMO OFDM Systems: A Reinforcement Learning Approach 25 Jan 2021 · 0 repositories · arXiv:2101.10300
-
Fire Threat Detection From Videos with Q-Rough Sets 21 Jan 2021 · 0 repositories · arXiv:2101.08459
-
Benchmarking Perturbation-based Saliency Maps for Explaining Atari Agents 18 Jan 2021 · 1 repository · arXiv:2101.07312
-
Randomized Ensembled Double Q-Learning: Learning Fast Without a Model 15 Jan 2021 · 6 repositories · arXiv:2101.05982Syntology community repositories only · 4 ran (of which 4 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified; every one of the 4 samples that ran constructed an object rather than computing a result (of 5 harvested samples) · 3 pointer-only (licence)
-
Continuous Deep Q-Learning with Simulator for Stabilization of Uncertain Discrete-Time Systems 13 Jan 2021 · 1 repository · arXiv:2101.05640
-
Action Priors for Large Action Spaces in Robotics 11 Jan 2021 · 1 repository · arXiv:2101.04178
-
Learning Augmented Index Policy for Optimal Service Placement at the Network Edge 10 Jan 2021 · 0 repositories · arXiv:2101.03641
-
Robust and Scalable Routing with Multi-Agent Deep Reinforcement Learning for MANETs 9 Jan 2021 · 0 repositories · arXiv:2101.03273
-
Evolving Reinforcement Learning Algorithms 8 Jan 2021 · 5 repositories · arXiv:2101.03958
-
Federated Intelligence for Active Queue Management in Inter-Domain Congestion 8 Jan 2021 · 1 repository
-
Reinforced Imitative Graph Representation Learning for Mobile User Profiling: An Adversarial Training Perspective 7 Jan 2021 · 0 repositories · arXiv:2101.02634
-
Reinforcement Learning with Latent Flow 6 Jan 2021 · 2 repositories · arXiv:2101.01857Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
A novel policy for pre-trained Deep Reinforcement Learning for Speech Emotion Recognition 4 Jan 2021 · 1 repository · arXiv:2101.00738
-
Addressing Distribution Shift in Online Reinforcement Learning with Offline Datasets 1 Jan 2021 · 0 repositories
-
Deep Reinforcement Learning-based Anti-jamming Power Allocation in a Two-cell NOMA Network 1 Jan 2021 · 0 repositories · arXiv:2101.00270
-
Double Q-learning: New Analysis and Sharper Finite-time Bound 1 Jan 2021 · 0 repositories
-
Learning Movement Strategies for Moving Target Defense 1 Jan 2021 · 0 repositories
-
Learning to Search for Fast Maximum Common Subgraph Detection 1 Jan 2021 · 0 repositories
-
PAC-Bayesian Randomized Value Function with Informative Prior 1 Jan 2021 · 0 repositories
-
Preventing Value Function Collapse in Ensemble Q-Learning by Maximizing Representation Diversity 1 Jan 2021 · 0 repositories
-
Uncertainty Weighted Offline Reinforcement Learning 1 Jan 2021 · 0 repositories
-
Weighted Bellman Backups for Improved Signal-to-Noise in Q-Updates 1 Jan 2021 · 0 repositories
-
Disentangled Planning and Control in Vision Based Robotics via Reward Machines 28 Dec 2020 · 0 repositories · arXiv:2012.14464
-
POPO: Pessimistic Offline Policy Optimization 26 Dec 2020 · 1 repository · arXiv:2012.13682
-
A State Representation Dueling Network for Deep Reinforcement Learning 24 Dec 2020 · 0 repositories
-
Assured RL: Reinforcement Learning with Almost Sure Constraints 24 Dec 2020 · 0 repositories · arXiv:2012.13036
-
Distributed Q-Learning with State Tracking for Multi-agent Networked Control 22 Dec 2020 · 0 repositories · arXiv:2012.12383
-
Goal Reasoning by Selecting Subgoals with Deep Q-Learning 22 Dec 2020 · 0 repositories · arXiv:2012.12335
-
Deploying Reinforcement Learning in Water Transport 14 Dec 2020 · 0 repositories
-
Learn to Play Tetris with Deep Reinforcement Learning 14 Dec 2020 · 0 repositories
-
Mobile Robots Autonomous Exploration with Reinforcement Learning 14 Dec 2020 · 0 repositories
-
Policy Gradient RL Algorithms as Directed Acyclic Graphs 14 Dec 2020 · 1 repository · arXiv:2012.07763
-
Semi-Supervised Off Policy Reinforcement Learning 9 Dec 2020 · 0 repositories · arXiv:2012.04809
-
Selective Pseudo-Labeling with Reinforcement Learning for Semi-Supervised Domain Adaptation 7 Dec 2020 · 0 repositories · arXiv:2012.03438
-
Hippocampal representations emerge when training recurrent neural networks on a memory dependent maze navigation task 2 Dec 2020 · 0 repositories · arXiv:2012.01328
-
Self-correcting Q-Learning 2 Dec 2020 · 0 repositories · arXiv:2012.01100
-
A new convergent variant of Q-learning with linear function approximation 1 Dec 2020 · 0 repositories
-
A Unified Switching System Perspective and Convergence Analysis of Q-Learning Algorithms 1 Dec 2020 · 0 repositories
-
Can Temporal-Difference and Q-Learning Learn Representation? A Mean-Field Theory 1 Dec 2020 · 0 repositories
-
Robust Multi-Agent Reinforcement Learning with Model Uncertainty 1 Dec 2020 · 0 repositories