Methods › Reinforcement Learning › Off-Policy TD Control › Q-Learning › Papers, page 2
Q-Learning
Papers archive 2025-07-28
archive papers tagged: 1,734 · with a code link: 464 · where Syntology ran a sample: 126 (105 with a run with no instrument failure, 21 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (126 of 1,734 tagged: 105 with a run with no instrument failure, 21 where every run was a failure of Syntology's instrument)
Page 2 of 18: papers 101 to 200 of 1,734, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Projection Implicit Q-Learning with Support Constraint for Offline Reinforcement Learning 15 Jan 2025 · 0 repositories · arXiv:2501.08907
-
Online inductive learning from answer sets for efficient reinforcement learning exploration 13 Jan 2025 · 0 repositories · arXiv:2501.07445
-
Perception-Guided EEG Analysis: A Deep Learning Approach Inspired by Level of Detail (LOD) Theory 11 Jan 2025 · 0 repositories · arXiv:2501.10428
-
Session-Level Dynamic Ad Load Optimization using Offline Robust Reinforcement Learning 9 Jan 2025 · 0 repositories · arXiv:2501.05591
-
β-DQN: Improving Deep Q-Learning By Evolving the Behavior 1 Jan 2025 · 0 repositories · arXiv:2501.00913
-
REM: A Scalable Reinforced Multi-Expert Framework for Multiplex Influence Maximization 1 Jan 2025 · 0 repositories · arXiv:2501.00779
-
Dynamic Optimization of Storage Systems Using Reinforcement Learning Techniques 29 Dec 2024 · 0 repositories · arXiv:2501.00068
-
Protein Structure Prediction in the 3D HP Model Using Deep Reinforcement Learning 29 Dec 2024 · 0 repositories · arXiv:2412.20329
-
MobileNetV2: A lightweight classification model for home-based sleep apnea screening 28 Dec 2024 · 1 repository · arXiv:2412.19967
-
A Reinforcement Learning-Based Task Mapping Method to Improve the Reliability of Clustered Manycores 26 Dec 2024 · 0 repositories · arXiv:2412.19340
-
Deep Learning-Based Traffic-Aware Base Station Sleep Mode and Cell Zooming Strategy in RIS-Aided Multi-Cell Networks 25 Dec 2024 · 0 repositories · arXiv:2412.18983
-
HyperQ-Opt: Q-learning for Hyperparameter Optimization 23 Dec 2024 · 0 repositories · arXiv:2412.17765
-
ACL-QL: Adaptive Conservative Level in Q-Learning for Offline Reinforcement Learning 22 Dec 2024 · 0 repositories · arXiv:2412.16848
-
Multi-Agent Q-Learning for Real-Time Load Balancing User Association and Handover in Mobile Networks 22 Dec 2024 · 0 repositories · arXiv:2412.19835
-
On Enhancing Network Throughput using Reinforcement Learning in Sliced Testbeds 21 Dec 2024 · 0 repositories · arXiv:2412.16673
-
Decoding fairness: a reinforcement learning perspective 20 Dec 2024 · 1 repository · arXiv:2412.16249
-
Distribution-Free Uncertainty Quantification in Mechanical Ventilation Treatment: A Conformal Deep Q-Learning Framework 17 Dec 2024 · 0 repositories · arXiv:2412.12597
-
Neural-Network-Driven Reward Prediction as a Heuristic: Advancing Q-Learning for Mobile Robot Path Planning 17 Dec 2024 · 0 repositories · arXiv:2412.12650
-
Integrated trucks assignment and scheduling problem with mixed service mode docks: A Q-learning based adaptive large neighborhood search algorithm 12 Dec 2024 · 0 repositories · arXiv:2412.09090
-
PickLLM: Context-Aware RL-Assisted Large Language Model Routing 12 Dec 2024 · 0 repositories · arXiv:2412.12170
-
Edge Delayed Deep Deterministic Policy Gradient: efficient continuous control for edge scenarios 9 Dec 2024 · 0 repositories · arXiv:2412.06390
-
DRL4AOI: A DRL Framework for Semantic-aware AOI Segmentation in Location-Based Services 6 Dec 2024 · 1 repository · arXiv:2412.05437
-
Demonstration Selection for In-Context Learning via Reinforcement Learning 5 Dec 2024 · 0 repositories · arXiv:2412.03966
-
Comparative Analysis of Multi-Agent Reinforcement Learning Policies for Crop Planning Decision Support 3 Dec 2024 · 0 repositories · arXiv:2412.02057
-
ECG-SleepNet: Deep Learning-Based Comprehensive Sleep Stage Classification Using ECG Signals 2 Dec 2024 · 0 repositories · arXiv:2412.01929
-
Q-learning-based Model-free Safety Filter 29 Nov 2024 · 0 repositories · arXiv:2411.19809
-
Dynamic Retail Pricing via Q-Learning -- A Reinforcement Learning Framework for Enhanced Revenue Management 27 Nov 2024 · 0 repositories · arXiv:2411.18261
-
Pretrained LLM Adapted with LoRA as a Decision Transformer for Offline RL in Quantitative Trading 26 Nov 2024 · 1 repository · arXiv:2411.17900
-
Time-Scale Separation in Q-Learning: Extending TD(△) for Action-Value Function Decomposition 21 Nov 2024 · 0 repositories · arXiv:2411.14019
-
Structure learning with Temporal Gaussian Mixture for model-based Reinforcement Learning 18 Nov 2024 · 0 repositories · arXiv:2411.11511
-
Mitigating Relative Over-Generalization in Multi-Agent Reinforcement Learning 17 Nov 2024 · 0 repositories · arXiv:2411.11099
-
Reinforcing Competitive Multi-Agents for Playing So Long Sucker 17 Nov 2024 · 0 repositories · arXiv:2411.11057
-
Innate-Values-driven Reinforcement Learning based Cognitive Modeling 14 Nov 2024 · 0 repositories · arXiv:2411.09160
-
Coverage Analysis for Digital Cousin Selection -- Improving Multi-Environment Q-Learning 13 Nov 2024 · 0 repositories · arXiv:2411.08360
-
Enhancing Robot Assistive Behaviour with Reinforcement Learning and Theory of Mind 11 Nov 2024 · 1 repository · arXiv:2411.07003
-
Q-SFT: Q-Learning for Language Models via Supervised Fine-Tuning 7 Nov 2024 · 0 repositories · arXiv:2411.05193
-
Think Smart, Act SMARL! Analyzing Probabilistic Logic Shields for Multi-Agent Reinforcement Learning 7 Nov 2024 · 1 repository · arXiv:2411.04867
-
A Comparative Study of Deep Reinforcement Learning for Crop Production Management 6 Nov 2024 · 0 repositories · arXiv:2411.04106
-
Beyond The Rainbow: High Performance Deep Reinforcement Learning on a Desktop PC 6 Nov 2024 · 3 repositories · arXiv:2411.03820Syntology official (archive's flag): 3 ran · 16 ran (of which 10 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 0 violated, 13 with no contract checked; 3 where Syntology's instrument failed) · 9 unverified (of 25 harvested samples) · 21 pointer-only (licence)
-
Temporal-Difference Learning Using Distributed Error Signals 6 Nov 2024 · 1 repository · arXiv:2411.03604Syntology official (archive's flag): 2 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 1 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 14 unverified (of 24 harvested samples) · 1 pointer-only (licence)
-
Dynamic Weight Adjusting Deep Q-Networks for Real-Time Environmental Adaptation 4 Nov 2024 · 0 repositories · arXiv:2411.02559
-
Simulation of Nanorobots with Artificial Intelligence and Reinforcement Learning for Advanced Cancer Cell Detection and Tracking 4 Nov 2024 · 1 repository · arXiv:2411.02345
-
HAVER: Instance-Dependent Error Bounds for Maximum Mean Estimation and Applications to Q-Learning and Monte Carlo Tree Search 1 Nov 2024 · 0 repositories · arXiv:2411.00405
-
CALE: Continuous Arcade Learning Environment 31 Oct 2024 · 1 repository · arXiv:2410.23810Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Q-learning for Quantile MDPs: A Decomposition, Performance, and Convergence Analysis 31 Oct 2024 · 1 repository · arXiv:2410.24128
-
Zonal RL-RRT: Integrated RL-RRT Path Planning with Collision Probability and Zone Connectivity 31 Oct 2024 · 1 repository · arXiv:2410.24205
-
ReferEverything: Towards Segmenting Everything We Can Speak of in Videos 30 Oct 2024 · 0 repositories · arXiv:2410.23287
-
Self-Driving Car Racing: Application of Deep Reinforcement Learning 30 Oct 2024 · 0 repositories · arXiv:2410.22766
-
Q-Distribution guided Q-learning for offline reinforcement learning: Uncertainty penalized Q-value via consistency model 27 Oct 2024 · 1 repository · arXiv:2410.20312Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Optimizing Load Scheduling in Power Grids Using Reinforcement Learning and Markov Decision Processes 23 Oct 2024 · 0 repositories · arXiv:2410.17696
-
A Novel Reinforcement Learning Model for Post-Incident Malware Investigations 19 Oct 2024 · 0 repositories · arXiv:2410.15028
-
Streaming Deep Reinforcement Learning Finally Works 18 Oct 2024 · 1 repository · arXiv:2410.14606Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 1 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Spectrum Sharing using Deep Reinforcement Learning in Vehicular Networks 16 Oct 2024 · 0 repositories · arXiv:2410.12521
-
DIAR: Diffusion-model-guided Implicit Q-learning with Adaptive Revaluation 15 Oct 2024 · 0 repositories · arXiv:2410.11338
-
Diffusion-Based Offline RL for Improved Decision-Making in Augmented ARC Task 15 Oct 2024 · 0 repositories · arXiv:2410.11324
-
Learning Agents With Prioritization and Parameter Noise in Continuous State and Action Space 15 Oct 2024 · 0 repositories · arXiv:2410.11250
-
Multi-Objective-Optimization Multi-AUV Assisted Data Collection Framework for IoUT Based on Offline Reinforcement Learning 15 Oct 2024 · 0 repositories · arXiv:2410.11282
-
Asymptotic Analysis of Sample-averaged Q-learning 14 Oct 2024 · 0 repositories · arXiv:2410.10737
-
Online waveform selection for cognitive radar 14 Oct 2024 · 0 repositories · arXiv:2410.10591
-
Hybrid LLM-DDQN based Joint Optimization of V2I Communication and Autonomous Driving 11 Oct 2024 · 0 repositories · arXiv:2410.08854
-
Gap-Dependent Bounds for Q-Learning using Reference-Advantage Decomposition 10 Oct 2024 · 0 repositories · arXiv:2410.07574
-
UNIQ: Offline Inverse Q-learning for Avoiding Undesirable Demonstrations 10 Oct 2024 · 0 repositories · arXiv:2410.08307
-
VerifierQ: Enhancing LLM Test Time Compute with Q-Learning-based Verifiers 10 Oct 2024 · 0 repositories · arXiv:2410.08048
-
Q-WSL: Optimizing Goal-Conditioned RL with Weighted Supervised Learning via Dynamic Programming 9 Oct 2024 · 0 repositories · arXiv:2410.06648
-
Learning in complex action spaces without policy gradients 8 Oct 2024 · 0 repositories · arXiv:2410.06317
-
Mimicking Human Intuition: Cognitive Belief-Driven Q-Learning 2 Oct 2024 · 0 repositories · arXiv:2410.01739
-
Efficient and Private Marginal Reconstruction with Local Non-Negativity 1 Oct 2024 · 1 repository · arXiv:2410.01091Syntology official (archive's flag): 17 ran · 17 ran (of which 2 constructed an object rather than computing a result; 12 with no instrument failure: 10 honoured, 0 violated, 2 with no contract checked; 5 where Syntology's instrument failed) · 8 unverified (of 25 harvested samples) · 25 pointer-only (licence)
-
Optimizing Photoplethysmography-Based Sleep Staging Models by Leveraging Temporal Context for Wearable Devices Applications 1 Oct 2024 · 0 repositories · arXiv:2410.00693
-
Adaptive Knowledge-based Multi-Objective Evolutionary Algorithm for Hybrid Flow Shop Scheduling Problems with Multiple Parallel Batch Processing Stages 27 Sep 2024 · 0 repositories · arXiv:2409.18524
-
Experimental Evaluation of Machine Learning Models for Goal-oriented Customer Service Chatbot with Pipeline Architecture 27 Sep 2024 · 0 repositories · arXiv:2409.18568
-
Optimized Monte Carlo Tree Search for Enhanced Decision Making in the FrozenLake Environment 25 Sep 2024 · 0 repositories · arXiv:2409.16620
-
A Multi-Agent Multi-Environment Mixed Q-Learning for Partially Decentralized Wireless Network Optimization 24 Sep 2024 · 1 repository · arXiv:2409.16450
-
Agent-state based policies in POMDPs: Beyond belief-state MDPs 24 Sep 2024 · 0 repositories · arXiv:2409.15703
-
Learning to Play Video Games with Intuitive Physics Priors 20 Sep 2024 · 0 repositories · arXiv:2409.13886
-
Data-Efficient Quadratic Q-Learning Using LMIs 18 Sep 2024 · 0 repositories · arXiv:2409.11986
-
Automating proton PBS treatment planning for head and neck cancers using policy gradient-based deep reinforcement learning 17 Sep 2024 · 0 repositories · arXiv:2409.11576
-
Audio-Driven Reinforcement Learning for Head-Orientation in Naturalistic Environments 16 Sep 2024 · 1 repository · arXiv:2409.10048
-
Offline Reinforcement Learning for Learning to Dispatch for Job Shop Scheduling 16 Sep 2024 · 1 repository · arXiv:2409.10589
-
KAN v.s. MLP for Offline Reinforcement Learning 15 Sep 2024 · 0 repositories · arXiv:2409.09653
-
Learning Discrete World Models for Heuristic Search 14 Sep 2024 · 1 repository
-
Deep reinforcement learning for tracking a moving target in jellyfish-like swimming 13 Sep 2024 · 0 repositories · arXiv:2409.08815
-
Autonomous Vehicle Decision-Making Framework for Considering Malicious Behavior at Unsignalized Intersections 11 Sep 2024 · 0 repositories · arXiv:2409.17162
-
Double Successive Over-Relaxation Q-Learning with an Extension to Deep Reinforcement Learning 10 Sep 2024 · 1 repository · arXiv:2409.06356
-
Reinforcement Learning for Rate Maximization in IRS-aided OWC Networks 7 Sep 2024 · 0 repositories · arXiv:2409.04842
-
Reward-Directed Score-Based Diffusion Models via q-Learning 7 Sep 2024 · 0 repositories · arXiv:2409.04832
-
Faster Q-Learning Algorithms for Restless Bandits 6 Sep 2024 · 0 repositories · arXiv:2409.05908
-
Whittle Index Learning Algorithms for Restless Bandits with Constant Stepsizes 6 Sep 2024 · 0 repositories · arXiv:2409.04605
-
Robust Q-Learning under Corrupted Rewards 5 Sep 2024 · 1 repository · arXiv:2409.03237
-
Reinforcement Learning-enabled Satellite Constellation Reconfiguration and Retasking for Mission-Critical Applications 3 Sep 2024 · 0 repositories · arXiv:2409.02270
-
Accelerated Multi-objective Task Learning using Modified Q-learning Algorithm 2 Sep 2024 · 0 repositories · arXiv:2409.01046
-
Stability of multiplexed NCS based on an epsilon-greedy algorithm for communication selection 2 Sep 2024 · 0 repositories · arXiv:2409.00949
-
The Sample-Communication Complexity Trade-off in Federated Q-Learning 30 Aug 2024 · 0 repositories · arXiv:2408.16981
-
Coverage Analysis of Multi-Environment Q-Learning Algorithms for Wireless Network Optimization 29 Aug 2024 · 0 repositories · arXiv:2408.16882
-
On Convergence of Average-Reward Q-Learning in Weakly Communicating Markov Decision Processes 29 Aug 2024 · 0 repositories · arXiv:2408.16262
-
Optimizing TD3 for 7-DOF Robotic Arm Grasping: Overcoming Suboptimality with Exploration-Enhanced Contrastive Learning 26 Aug 2024 · 0 repositories · arXiv:2408.14009
-
Can LLM be a Good Path Planner based on Prompt Engineering? Mitigating the Hallucination for Path Planning 23 Aug 2024 · 0 repositories · arXiv:2408.13184
-
Deviations from the Nash equilibrium and emergence of tacit collusion in a two-player optimal execution game with reinforcement learning 21 Aug 2024 · 0 repositories · arXiv:2408.11773
-
Efficient Exploration in Deep Reinforcement Learning: A Novel Bayesian Actor-Critic Algorithm 19 Aug 2024 · 0 repositories · arXiv:2408.10055
-
A Conflicts-free, Speed-lossless KAN-based Reinforcement Learning Decision System for Interactive Driving in Roundabouts 15 Aug 2024 · 0 repositories · arXiv:2408.08242
-
Explaining an Agent's Future Beliefs through Temporally Decomposing Future Reward Estimators 15 Aug 2024 · 1 repository · arXiv:2408.08230