Methods › Reinforcement Learning › Off-Policy TD Control › Q-Learning › Papers, page 7
Q-Learning
Papers archive 2025-07-28
archive papers tagged: 1,734 · with a code link: 464 · where Syntology ran a sample: 126 (105 with a run with no instrument failure, 21 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (126 of 1,734 tagged: 105 with a run with no instrument failure, 21 where every run was a failure of Syntology's instrument)
Page 7 of 18: papers 601 to 700 of 1,734, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
MACOptions: Multi-Agent Learning with Centralized Controller and Options Framework 7 Feb 2023 · 0 repositories · arXiv:2302.03800
-
Deep Reinforcement Learning for Traffic Light Control in Intelligent Transportation Systems 4 Feb 2023 · 0 repositories · arXiv:2302.03669
-
Accelerating Policy Gradient by Estimating Value Function from Prior Computation in Deep Reinforcement Learning 2 Feb 2023 · 0 repositories · arXiv:2302.01399
-
Best Possible Q-Learning 2 Feb 2023 · 0 repositories · arXiv:2302.01188
-
Diversity Through Exclusion (DTE): Niche Identification for Reinforcement Learning through Value-Decomposition 2 Feb 2023 · 0 repositories · arXiv:2302.01180
-
Sample Complexity of Kernel-Based Q-Learning 1 Feb 2023 · 0 repositories · arXiv:2302.00727
-
Sample Efficient Deep Reinforcement Learning via Local Planning 29 Jan 2023 · 0 repositories · arXiv:2301.12579
-
Analyzing Robustness of the Deep Reinforcement Learning Algorithm in Ramp Metering Applications Considering False Data Injection Attack and Defense 28 Jan 2023 · 0 repositories · arXiv:2301.12036
-
The impact of surplus sharing on the outcomes of specific investments under negotiated transfer pricing: An agent-based simulation with fuzzy Q-learning agents 28 Jan 2023 · 0 repositories · arXiv:2301.12255
-
Single-Trajectory Distributionally Robust Reinforcement Learning 27 Jan 2023 · 0 repositories · arXiv:2301.11721
-
FedHQL: Federated Heterogeneous Q-Learning 26 Jan 2023 · 0 repositories · arXiv:2301.11135
-
Learning from Multiple Independent Advisors in Multi-agent Reinforcement Learning 26 Jan 2023 · 1 repository · arXiv:2301.11153
-
Asymptotic Convergence and Performance of Multi-Agent Q-Learning Dynamics 23 Jan 2023 · 0 repositories · arXiv:2301.09619
-
Asynchronous Deep Double Duelling Q-Learning for Trading-Signal Execution in Limit Order Book Markets 20 Jan 2023 · 0 repositories · arXiv:2301.08688
-
Modeling Moral Choices in Social Dilemmas with Multi-Agent Reinforcement Learning 20 Jan 2023 · 2 repositories · arXiv:2301.08491Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
Risk-Averse Reinforcement Learning via Dynamic Time-Consistent Risk Measures 14 Jan 2023 · 0 repositories · arXiv:2301.05981
-
Decentralized model-free reinforcement learning in stochastic games with average-reward objective 13 Jan 2023 · 0 repositories · arXiv:2301.05630
-
Hierarchical Deep Q-Learning Based Handover in Wireless Networks with Dual Connectivity 13 Jan 2023 · 0 repositories · arXiv:2301.05391
-
TransfQMix: Transformers for Leveraging the Graph Structure of Multi-Agent Reinforcement Learning Problems 13 Jan 2023 · 1 repository · arXiv:2301.05334
-
schlably: A Python Framework for Deep Reinforcement Learning Based Scheduling Experiments 10 Jan 2023 · 1 repository · arXiv:2301.04182
-
Tuning Path Tracking Controllers for Autonomous Cars Using Reinforcement Learning 9 Jan 2023 · 0 repositories · arXiv:2301.03363
-
XDQN: Inherently Interpretable DQN through Mimicking 8 Jan 2023 · 0 repositories · arXiv:2301.03043
-
Extreme Q-Learning: MaxEnt RL without Entropy 5 Jan 2023 · 4 repositories · arXiv:2301.02328Syntology 8 ran (of which 5 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 3 where Syntology's instrument failed) · 5 unverified (of 13 harvested samples) · 7 pointer-only (licence)
-
Learning a Generic Value-Selection Heuristic Inside a Constraint Programming Solver 5 Jan 2023 · 1 repository · arXiv:2301.01913
-
Deep Spectral Q-learning with Application to Mobile Health 3 Jan 2023 · 0 repositories · arXiv:2301.00927
-
Hierarchical Deep Reinforcement Learning for Age-of-Information Minimization in IRS-aided and Wireless-powered Wireless Networks 27 Dec 2022 · 0 repositories · arXiv:2212.13390
-
Decoding surface codes with deep reinforcement learning and probabilistic policy reuse 22 Dec 2022 · 0 repositories · arXiv:2212.11890
-
Control of Continuous Quantum Systems with Many Degrees of Freedom based on Convergent Reinforcement Learning 21 Dec 2022 · 1 repository · arXiv:2212.10705
-
Neighboring state-based RL Exploration 21 Dec 2022 · 0 repositories · arXiv:2212.10712
-
Taming Lagrangian Chaos with Multi-Objective Reinforcement Learning 19 Dec 2022 · 0 repositories · arXiv:2212.09612
-
Offline Robot Reinforcement Learning with Uncertainty-Guided Human Expert Sampling 16 Dec 2022 · 0 repositories · arXiv:2212.08232
-
Frugal Reinforcement-based Active Learning 9 Dec 2022 · 0 repositories · arXiv:2212.04868
-
PALMER: Perception-Action Loop with Memory for Long-Horizon Planning 8 Dec 2022 · 0 repositories · arXiv:2212.04581
-
EASpace: Enhanced Action Space for Policy Transfer 7 Dec 2022 · 1 repository · arXiv:2212.03540
-
A Machine with Short-Term, Episodic, and Semantic Memory Systems 5 Dec 2022 · 1 repository · arXiv:2212.02098Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 10 harvested samples)
-
Automata Learning meets Shielding 4 Dec 2022 · 1 repository · arXiv:2212.01838
-
Automatic Discovery of Multi-perspective Process Model using Reinforcement Learning 30 Nov 2022 · 0 repositories · arXiv:2211.16687
-
Welfare and Fairness in Multi-objective Reinforcement Learning 30 Nov 2022 · 1 repository · arXiv:2212.01382Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
ACE: Cooperative Multi-agent Q-learning with Bidirectional Action-Dependency 29 Nov 2022 · 1 repository · arXiv:2211.16068Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
Airfoil Shape Optimization using Deep Q-Network 29 Nov 2022 · 0 repositories · arXiv:2211.17189
-
Offline Q-Learning on Diverse Multi-Task Data Both Scales And Generalizes 28 Nov 2022 · 0 repositories · arXiv:2211.15144
-
QLAMMP: A Q-Learning Agent for Optimizing Fees on Automated Market Making Protocols 28 Nov 2022 · 0 repositories · arXiv:2211.14977
-
Applying Deep Reinforcement Learning to the HP Model for Protein Structure Prediction 27 Nov 2022 · 1 repository · arXiv:2211.14939Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 13 harvested samples)
-
Reinforcement Causal Structure Learning on Order Graph 22 Nov 2022 · 0 repositories · arXiv:2211.12151
-
Examining Policy Entropy of Reinforcement Learning Agents for Personalization Tasks 21 Nov 2022 · 1 repository · arXiv:2211.11869
-
Simultaneously Updating All Persistence Values in Reinforcement Learning 21 Nov 2022 · 0 repositories · arXiv:2211.11620
-
Analysis of Reinforcement Learning Schemes for Trajectory Optimization of an Aerial Radio Unit 18 Nov 2022 · 0 repositories · arXiv:2211.10524
-
Credit-cognisant reinforcement learning for multi-agent cooperation 18 Nov 2022 · 0 repositories · arXiv:2211.10100
-
A Reinforcement Learning Approach for Process Parameter Optimization in Additive Manufacturing 17 Nov 2022 · 0 repositories · arXiv:2211.09545
-
Planning Irregular Object Packing via Hierarchical Reinforcement Learning 17 Nov 2022 · 0 repositories · arXiv:2211.09382
-
Solar Power driven EV Charging Optimization with Deep Reinforcement Learning 17 Nov 2022 · 0 repositories · arXiv:2211.09479
-
Addressing the issue of stochastic environments and local decision-making in multi-objective reinforcement learning 16 Nov 2022 · 0 repositories · arXiv:2211.08669
-
On the Global Convergence of Fitted Q-Iteration with Two-layer Neural Network Parametrization 14 Nov 2022 · 0 repositories · arXiv:2211.07675
-
NREM and REM: cognitive and energetic gains in thalamo-cortical sleeping and awake spiking model 13 Nov 2022 · 0 repositories · arXiv:2211.06889
-
Deep W-Networks: Solving Multi-Objective Optimisation Problems With Deep Reinforcement Learning 9 Nov 2022 · 1 repository · arXiv:2211.04813
-
Reinforcement Learning in Non-Markovian Environments 3 Nov 2022 · 0 repositories · arXiv:2211.01595
-
Deep Reinforcement Learning for Power Control in Next-Generation WiFi Network Systems 2 Nov 2022 · 0 repositories · arXiv:2211.01107
-
DynamicLight: Two-Stage Dynamic Traffic Signal Timing 2 Nov 2022 · 1 repository · arXiv:2211.01025
-
Offline RL With Realistic Datasets: Heteroskedasticity and Support Constraints 2 Nov 2022 · 0 repositories · arXiv:2211.01052
-
Representation Learning for General-sum Low-rank Markov Games 30 Oct 2022 · 0 repositories · arXiv:2210.16976
-
DeFIX: Detecting and Fixing Failure Scenarios with Reinforcement Learning in Imitation Learning Based Autonomous Driving 29 Oct 2022 · 2 repositories · arXiv:2210.16567
-
Solving Continuous Control via Q-learning 22 Oct 2022 · 1 repository · arXiv:2210.12566Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 7 harvested samples)
-
Sufficient Exploration for Convex Q-learning 17 Oct 2022 · 0 repositories · arXiv:2210.09409
-
Model-Free Characterizations of the Hamilton-Jacobi-Bellman Equation and Convex Q-Learning in Continuous Time 14 Oct 2022 · 0 repositories · arXiv:2210.08131
-
Deep reinforcement learning for automatic run-time adaptation of UWB PHY radio settings 13 Oct 2022 · 0 repositories · arXiv:2210.15498
-
Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient 13 Oct 2022 · 1 repository · arXiv:2210.06718Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 3 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Sustainable Online Reinforcement Learning for Auto-bidding 13 Oct 2022 · 1 repository · arXiv:2210.07006
-
Censored Deep Reinforcement Patrolling with Information Criterion for Monitoring Large Water Resources using Autonomous Surface Vehicles 12 Oct 2022 · 0 repositories · arXiv:2210.08115
-
Factors of Influence of the Overestimation Bias of Q-Learning 11 Oct 2022 · 1 repository · arXiv:2210.05262
-
Pre-Training for Robots: Offline RL Enables Learning New Tasks from a Handful of Trials 11 Oct 2022 · 1 repository · arXiv:2210.05178
-
Elastic Step DQN: A novel multi-step algorithm to alleviate overestimation in Deep QNetworks 7 Oct 2022 · 0 repositories · arXiv:2210.03325
-
Reinforcement Learning Approach for Multi-Agent Flexible Scheduling Problems 7 Oct 2022 · 0 repositories · arXiv:2210.03674
-
Exploration Policies for On-the-Fly Controller Synthesis: A Reinforcement Learning Approach 7 Oct 2022 · 1 repository · arXiv:2210.05393
-
Towards Safe Mechanical Ventilation Treatment Using Deep Offline Reinforcement Learning 5 Oct 2022 · 1 repository · arXiv:2210.02552Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples)
-
Interpretable Option Discovery using Deep Q-Learning and Variational Autoencoders 3 Oct 2022 · 0 repositories · arXiv:2210.01231
-
Offline Reinforcement Learning with Differentiable Function Approximation is Provably Efficient 3 Oct 2022 · 0 repositories · arXiv:2210.00750
-
Bayesian Q-learning With Imperfect Expert Demonstrations 1 Oct 2022 · 0 repositories · arXiv:2210.01800
-
Deep Recurrent Q-learning for Energy-constrained Coverage with a Mobile Robot 1 Oct 2022 · 0 repositories · arXiv:2210.00327
-
On Convergence of Average-Reward Off-Policy Control Algorithms in Weakly Communicating MDPs 30 Sep 2022 · 0 repositories · arXiv:2209.15141
-
FIRE: A Failure-Adaptive Reinforcement Learning Framework for Edge Computing Migrations 28 Sep 2022 · 0 repositories · arXiv:2209.14399
-
Predictive Crypto-Asset Automated Market Making Architecture for Decentralized Finance using Deep Reinforcement Learning 28 Sep 2022 · 0 repositories · arXiv:2211.01346
-
Understanding Hindsight Goal Relabeling from a Divergence Minimization Perspective 26 Sep 2022 · 0 repositories · arXiv:2209.13046
-
Revisiting Discrete Soft Actor-Critic 21 Sep 2022 · 1 repository · arXiv:2209.10081
-
Comparative Study of Q-Learning and NeuroEvolution of Augmenting Topologies for Self Driving Agents 19 Sep 2022 · 0 repositories · arXiv:2209.09007
-
MAN: Multi-Action Networks Learning 19 Sep 2022 · 1 repository · arXiv:2209.09329
-
MA2QL: A Minimalist Approach to Fully Decentralized Multi-Agent Reinforcement Learning 17 Sep 2022 · 0 repositories · arXiv:2209.08244
-
M²DQN: A Robust Method for Accelerating Deep Q-learning Network 16 Sep 2022 · 1 repository · arXiv:2209.07809
-
Reducing Variance in Temporal-Difference Value Estimation via Ensemble of Deep Networks 16 Sep 2022 · 1 repository · arXiv:2209.07670
-
Reinforcement Learning-Based Cooperative P2P Power Trading between DC Nanogrid Clusters with Wind and PV Energy Resources 16 Sep 2022 · 0 repositories · arXiv:2209.07744
-
Deep Reinforcement Learning for Task Offloading in UAV-Aided Smart Farm Networks 15 Sep 2022 · 0 repositories · arXiv:2209.07367
-
IoT-Aerial Base Station Task Offloading with Risk-Sensitive Reinforcement Learning for Smart Agriculture 15 Sep 2022 · 0 repositories · arXiv:2209.07382
-
Pathfinding in Random Partially Observable Environments with Vision-Informed Deep Reinforcement Learning 11 Sep 2022 · 0 repositories · arXiv:2209.04801
-
Structured Q-learning For Antibody Design 10 Sep 2022 · 0 repositories · arXiv:2209.04698
-
Continual learning benefits from multiple sleep mechanisms: NREM, REM, and Synaptic Downscaling 9 Sep 2022 · 0 repositories · arXiv:2209.05245
-
Double Q-Learning for Citizen Relocation During Natural Hazards 8 Sep 2022 · 0 repositories · arXiv:2209.03800
-
Q-learning Decision Transformer: Leveraging Dynamic Programming for Conditional Sequence Modelling in Offline RL 8 Sep 2022 · 1 repository · arXiv:2209.03993
-
Reward Delay Attacks on Deep Reinforcement Learning 8 Sep 2022 · 1 repository · arXiv:2209.03540
-
Distilling Deep RL Models Into Interpretable Neuro-Fuzzy Systems 7 Sep 2022 · 0 repositories · arXiv:2209.03357
-
SlateFree: a Model-Free Decomposition for Reinforcement Learning with Slate Actions 5 Sep 2022 · 0 repositories · arXiv:2209.01876
-
A Technique to Create Weaker Abstract Board Game Agents via Reinforcement Learning 1 Sep 2022 · 0 repositories · arXiv:2209.00711