Methods › Reinforcement Learning › Off-Policy TD Control › Q-Learning › Papers, page 3
Q-Learning
Papers archive 2025-07-28
archive papers tagged: 1,734 · with a code link: 464 · where Syntology ran a sample: 126 (105 with a run with no instrument failure, 21 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (126 of 1,734 tagged: 105 with a run with no instrument failure, 21 where every run was a failure of Syntology's instrument)
Page 3 of 18: papers 201 to 300 of 1,734, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Variance-Reduced Cascade Q-learning: Algorithms and Sample Complexity 13 Aug 2024 · 0 repositories · arXiv:2408.06544
-
Crowd Intelligence for Early Misinformation Prediction on Social Media 8 Aug 2024 · 1 repository · arXiv:2408.04463
-
Research on Autonomous Driving Decision-making Strategies based Deep Reinforcement Learning 6 Aug 2024 · 0 repositories · arXiv:2408.03084
-
QADQN: Quantum Attention Deep Q-Network for Financial Market Prediction 6 Aug 2024 · 0 repositories · arXiv:2408.03088
-
Model-free optimal controller for discrete-time Markovian jump linear systems: A Q-learning approach 6 Aug 2024 · 0 repositories · arXiv:2408.03077
-
Whittle's index-based age-of-information minimization in multi-energy harvesting source networks 5 Aug 2024 · 0 repositories · arXiv:2408.02570
-
Multi-Objective Deep Reinforcement Learning for Optimisation in Autonomous Systems 2 Aug 2024 · 0 repositories · arXiv:2408.01188
-
Multi-agent Assessment with QoS Enhancement for HD Map Updates in a Vehicular Network 31 Jul 2024 · 0 repositories · arXiv:2407.21460
-
Evolution of cooperation in the public goods game with Q-learning 29 Jul 2024 · 0 repositories · arXiv:2407.19851
-
Evolution of cooperation with Q-learning: the impact of information perception 29 Jul 2024 · 0 repositories · arXiv:2407.19634
-
Multi-Agent Deep Reinforcement Learning for Energy Efficient Multi-Hop STAR-RIS-Assisted Transmissions 26 Jul 2024 · 0 repositories · arXiv:2407.18627
-
Principal-Agent Reinforcement Learning: Orchestrating AI Agents with Contracts 25 Jul 2024 · 0 repositories · arXiv:2407.18074
-
In Search for Architectures and Loss Functions in Multi-Objective Reinforcement Learning 23 Jul 2024 · 0 repositories · arXiv:2407.16807
-
Evaluation of Reinforcement Learning for Autonomous Penetration Testing using A3C, Q-learning and DQN 22 Jul 2024 · 0 repositories · arXiv:2407.15656
-
MODRL-TA:A Multi-Objective Deep Reinforcement Learning Framework for Traffic Allocation in E-Commerce Search 22 Jul 2024 · 0 repositories · arXiv:2407.15476
-
Coverage-aware and Reinforcement Learning Using Multi-agent Approach for HD Map QoS in a Realistic Environment 19 Jul 2024 · 0 repositories · arXiv:2408.03329
-
An Agile Adaptation Method for Multi-mode Vehicle Communication Networks 18 Jul 2024 · 0 repositories · arXiv:2408.01429
-
Optimistic Q-learning for average reward and episodic reinforcement learning 18 Jul 2024 · 0 repositories · arXiv:2407.13743
-
Solving the Model Unavailable MARE using Q-Learning Algorithm 18 Jul 2024 · 0 repositories · arXiv:2407.13227
-
Cooperative Reward Shaping for Multi-Agent Pathfinding 15 Jul 2024 · 0 repositories · arXiv:2407.10403
-
Reinforcement Learning in High-frequency Market Making 14 Jul 2024 · 1 repository · arXiv:2407.21025
-
AlphaDou: High-Performance End-to-End Doudizhu AI Integrating Bidding 14 Jul 2024 · 1 repository · arXiv:2407.10279
-
PAIL: Performance based Adversarial Imitation Learning Engine for Carbon Neutral Optimization 12 Jul 2024 · 0 repositories · arXiv:2407.08910
-
PID Accelerated Temporal Difference Algorithms 11 Jul 2024 · 0 repositories · arXiv:2407.08803
-
Multi-agent Reinforcement Learning-based Network Intrusion Detection System 8 Jul 2024 · 0 repositories · arXiv:2407.05766
-
Periodic agent-state based Q-learning for POMDPs 8 Jul 2024 · 0 repositories · arXiv:2407.06121
-
A Multi-Step Minimax Q-learning Algorithm for Two-Player Zero-Sum Markov Games 5 Jul 2024 · 1 repository · arXiv:2407.04240Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Simplifying Deep Temporal Difference Learning 5 Jul 2024 · 1 repository · arXiv:2407.04811Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Unified continuous-time q-learning for mean-field game and mean-field control problems 5 Jul 2024 · 0 repositories · arXiv:2407.04521
-
Artificial Intelligence and Algorithmic Price Collusion in Two-sided Markets 4 Jul 2024 · 0 repositories · arXiv:2407.04088
-
Configuring Transmission Thresholds in IIoT Alarm Scenarios for Energy-Efficient Event Reporting 4 Jul 2024 · 0 repositories · arXiv:2407.03982
-
Continuous-time q-Learning for Jump-Diffusion Models under Tsallis Entropy 4 Jul 2024 · 0 repositories · arXiv:2407.03888
-
Q-Adapter: Customizing Pre-trained LLMs to New Preferences with Forgetting Mitigation 4 Jul 2024 · 1 repository · arXiv:2407.03856Syntology official (archive's flag): 2 ran · 13 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 4 where Syntology's instrument failed) · 1 unverified (of 14 harvested samples) · 1 pointer-only (licence)
-
Rethinking Data Augmentation for Robust LiDAR Semantic Segmentation in Adverse Weather 2 Jul 2024 · 2 repositories · arXiv:2407.02286
-
Two-Step Q-Learning 2 Jul 2024 · 0 repositories · arXiv:2407.02369
-
A Deep Reinforcement Learning Approach to Battery Management in Dairy Farming via Proximal Policy Optimization 1 Jul 2024 · 0 repositories · arXiv:2407.01653
-
Model-based Offline Reinforcement Learning with Lower Expectile Q-Learning 30 Jun 2024 · 0 repositories · arXiv:2407.00699
-
Contextualized Hybrid Ensemble Q-learning: Learning Fast with Control Priors 28 Jun 2024 · 1 repository · arXiv:2406.19768
-
Towards Secure and Efficient Data Scheduling for Vehicular Social Networks 28 Jun 2024 · 0 repositories · arXiv:2407.00141
-
MEReQ: Max-Ent Residual-Q Inverse RL for Sample-Efficient Alignment from Intervention 24 Jun 2024 · 0 repositories · arXiv:2406.16258
-
EduQate: Generating Adaptive Curricula through RMABs in Education Settings 20 Jun 2024 · 0 repositories · arXiv:2406.14122
-
Equivariant Offline Reinforcement Learning 20 Jun 2024 · 0 repositories · arXiv:2406.13961
-
Learning to Select Goals in Automated Planning with Deep-Q Learning 20 Jun 2024 · 0 repositories · arXiv:2406.14779
-
Reinforcement-Learning based routing for packet-optical networks with hybrid telemetry 18 Jun 2024 · 1 repository · arXiv:2406.12602
-
Catalytic evolution of cooperation in a population with behavioural bimodality 17 Jun 2024 · 0 repositories · arXiv:2406.11121
-
Optimal Transport-Assisted Risk-Sensitive Q-Learning 17 Jun 2024 · 0 repositories · arXiv:2406.11774
-
Post-hoc Utterance Refining Method by Entity Mining for Faithful Knowledge Grounded Conversations 16 Jun 2024 · 1 repository · arXiv:2406.10809
-
Mix Q-learning for Lane Changing: A Collaborative Decision-Making Method in Multi-Agent Deep Reinforcement Learning 14 Jun 2024 · 0 repositories · arXiv:2406.09755
-
Hadamard Representations: Augmenting Hyperbolic Tangents in RL 13 Jun 2024 · 0 repositories · arXiv:2406.09079
-
Multi-agent Reinforcement Learning with Deep Networks for Diverse Q-Vectors 12 Jun 2024 · 0 repositories · arXiv:2406.07848
-
Probing Implicit Bias in Semi-gradient Q-learning: Visualizing the Effective Loss Landscapes via the Fokker--Planck Equation 12 Jun 2024 · 1 repository · arXiv:2406.08148
-
PlanDQ: Hierarchical Plan Orchestration via D-Conductor and Q-Performer 10 Jun 2024 · 1 repository · arXiv:2406.06793
-
Online Policy Distillation with Decision-Attention 8 Jun 2024 · 0 repositories · arXiv:2406.05488
-
Fast-Fading Channel and Power Optimization of the Magnetic Inductive Cellular Network 7 Jun 2024 · 0 repositories · arXiv:2406.04737
-
Online Frequency Scheduling by Learning Parallel Actions 7 Jun 2024 · 0 repositories · arXiv:2406.05041
-
Stabilizing Extreme Q-learning by Maclaurin Expansion 7 Jun 2024 · 1 repository · arXiv:2406.04896
-
Bootstrapping Expectiles in Reinforcement Learning 6 Jun 2024 · 0 repositories · arXiv:2406.04081
-
Evaluating the Influence of Temporal Context on Automatic Mouse Sleep Staging through the Application of Human Models 6 Jun 2024 · 0 repositories · arXiv:2406.16911
-
Strategically Conservative Q-Learning 6 Jun 2024 · 1 repository · arXiv:2406.04534
-
Nonlinear Transformations Against Unlearnable Datasets 5 Jun 2024 · 0 repositories · arXiv:2406.02883
-
Algorithmic Collusion in Dynamic Pricing with Deep Reinforcement Learning 4 Jun 2024 · 0 repositories · arXiv:2406.02437
-
Towards Universal and Black-Box Query-Response Only Attack on LLMs with QROA 4 Jun 2024 · 1 repository · arXiv:2406.02044
-
Tabular and Deep Learning for the Whittle Index 4 Jun 2024 · 0 repositories · arXiv:2406.02057
-
A New View on Planning in Online Reinforcement Learning 3 Jun 2024 · 0 repositories · arXiv:2406.01562
-
How to discretize continuous state-action spaces in Q-learning: A symbolic control approach 3 Jun 2024 · 0 repositories · arXiv:2406.01548
-
Diffusion Policies creating a Trust Region for Offline Reinforcement Learning 30 May 2024 · 1 repository · arXiv:2405.19690Syntology official (archive's flag): 17 ran · 17 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 6 where Syntology's instrument failed) · 3 unverified (of 20 harvested samples) · 20 pointer-only (licence)
-
Q-learning as a monotone scheme 30 May 2024 · 0 repositories · arXiv:2405.20538
-
Federated Q-Learning with Reference-Advantage Decomposition: Almost Optimal Regret and Logarithmic Communication Cost 29 May 2024 · 0 repositories · arXiv:2405.18795
-
AlignIQL: Policy Alignment in Implicit Q-Learning through Constrained Optimization 28 May 2024 · 1 repository · arXiv:2405.18187
-
Mutation-Bias Learning in Games 28 May 2024 · 0 repositories · arXiv:2405.18190
-
Safe Multi-Agent Reinforcement Learning with Bilevel Optimization in Autonomous Driving 28 May 2024 · 1 repository · arXiv:2405.18209
-
Spatial-temporal analysis of neural desynchronization in sleep-like states reveals critical dynamics 28 May 2024 · 0 repositories · arXiv:2405.18329
-
Analysis of Multiscale Reinforcement Q-Learning Algorithms for Mean Field Control Games 27 May 2024 · 0 repositories · arXiv:2405.17017
-
Biological Neurons Compete with Deep Reinforcement Learning in Sample Efficiency in a Simulated Gameworld 27 May 2024 · 0 repositories · arXiv:2405.16946
-
An Evolutionary Framework for Connect-4 as Test-Bed for Comparison of Advanced Minimax, Q-Learning and MCTS 26 May 2024 · 0 repositories · arXiv:2405.16595
-
Reinforcement Learning for Jump-Diffusions, with Financial Applications 26 May 2024 · 0 repositories · arXiv:2405.16449
-
Extracting Heuristics from Large Language Models for Reward Shaping in Reinforcement Learning 24 May 2024 · 0 repositories · arXiv:2405.15194
-
Knowledge-Informed Auto-Penetration Testing Based on Reinforcement Learning with Reward Machine 24 May 2024 · 0 repositories · arXiv:2405.15908
-
A finite time analysis of distributed Q-learning 23 May 2024 · 0 repositories · arXiv:2405.14078
-
Exclusively Penalized Q-learning for Offline Reinforcement Learning 23 May 2024 · 0 repositories · arXiv:2405.14082
-
Traffic control using intelligent timing of traffic lights with reinforcement learning technique and real-time processing of surveillance camera images 22 May 2024 · 0 repositories · arXiv:2405.13256
-
Stochastic Q-learning for Large Discrete Action Spaces 16 May 2024 · 0 repositories · arXiv:2405.10310
-
Deep Reinforcement Learning for Real-Time Ground Delay Program Revision and Corresponding Flight Delay Assignments 14 May 2024 · 0 repositories · arXiv:2405.08298
-
An Initial Introduction to Cooperative Multi-Agent Reinforcement Learning 10 May 2024 · 0 repositories · arXiv:2405.06161
-
SwiftRL: Towards Efficient Reinforcement Learning on Real Processing-In-Memory Systems 7 May 2024 · 1 repository · arXiv:2405.03967
-
A Network Simulation of OTC Markets with Multiple Agents 3 May 2024 · 0 repositories · arXiv:2405.02480
-
Regularized Q-learning through Robust Averaging 3 May 2024 · 1 repository · arXiv:2405.02201
-
Zero-Sum Positional Differential Games as a Framework for Robust Reinforcement Learning: Deep Q-Learning Approach 3 May 2024 · 0 repositories · arXiv:2405.02044
-
LOQA: Learning with Opponent Q-Learning Awareness 2 May 2024 · 0 repositories · arXiv:2405.01035
-
Cell Switching in HAPS-Aided Networking: How the Obscurity of Traffic Loads Affects the Decision 1 May 2024 · 0 repositories · arXiv:2405.00387
-
Portfolio Management using Deep Reinforcement Learning 1 May 2024 · 0 repositories · arXiv:2405.01604
-
Numeric Reward Machines 30 Apr 2024 · 0 repositories · arXiv:2404.19370
-
Reinforcement Learning Problem Solving with Large Language Models 29 Apr 2024 · 0 repositories · arXiv:2404.18638
-
AFU: Actor-Free critic Updates in off-policy RL for continuous control 24 Apr 2024 · 1 repository · arXiv:2404.16159
-
Age of Information Minimization using Multi-agent UAVs based on AI-Enhanced Mean Field Resource Allocation 24 Apr 2024 · 0 repositories · arXiv:2405.00056
-
Recursive Backwards Q-Learning in Deterministic Environments 24 Apr 2024 · 0 repositories · arXiv:2404.15822
-
Research on Robot Path Planning Based on Reinforcement Learning 22 Apr 2024 · 1 repository · arXiv:2404.14077
-
Unified ODE Analysis of Smooth Q-Learning Algorithms 20 Apr 2024 · 0 repositories · arXiv:2404.14442
-
Continuous-time Risk-sensitive Reinforcement Learning via Quadratic Variation Penalty 19 Apr 2024 · 0 repositories · arXiv:2404.12598
-
Reinforcement Learning Approach for Integrating Compressed Contexts into Knowledge Graphs 19 Apr 2024 · 0 repositories · arXiv:2404.12587