Methods › Reinforcement Learning › Off-Policy TD Control › Q-Learning › Papers, page 8
Q-Learning
Papers archive 2025-07-28
archive papers tagged: 1,734 · with a code link: 464 · where Syntology ran a sample: 126 (105 with a run with no instrument failure, 21 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (126 of 1,734 tagged: 105 with a run with no instrument failure, 21 where every run was a failure of Syntology's instrument)
Page 8 of 18: papers 701 to 800 of 1,734, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Partial Counterfactual Identification for Infinite Horizon Partially Observable Markov Decision Process 31 Aug 2022 · 0 repositories · arXiv:2209.00137
-
Prospect Theory-inspired Automated P2P Energy Trading with Q-learning-based Dynamic Pricing 26 Aug 2022 · 0 repositories · arXiv:2208.12777
-
Recurrent Neural Network-based Anti-jamming Framework for Defense Against Multiple Jamming Policies 19 Aug 2022 · 0 repositories · arXiv:2208.09518
-
Diffusion Policies as an Expressive Policy Class for Offline Reinforcement Learning 12 Aug 2022 · 3 repositories · arXiv:2208.06193Syntology official (archive's flag): 1 ran · 11 ran (of which 5 constructed an object rather than computing a result; 11 with no instrument failure: 1 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 7 unverified (of 18 harvested samples) · 10 pointer-only (licence)
-
Disentangled Modeling of Domain and Relevance for Adaptable Dense Retrieval 11 Aug 2022 · 1 repository · arXiv:2208.05753
-
Prediction-based Hybrid Slicing Framework for Service Level Agreement Guarantee in Mobility Scenarios: A Deep Learning Approach 6 Aug 2022 · 0 repositories · arXiv:2208.03460
-
Reinforcement Learning for Joint V2I Network Selection and Autonomous Driving Policies 3 Aug 2022 · 0 repositories · arXiv:2208.02249
-
Digital Twin-Assisted Efficient Reinforcement Learning for Edge Task Scheduling 2 Aug 2022 · 0 repositories · arXiv:2208.01781
-
A Maintenance Planning Framework using Online and Offline Deep Reinforcement Learning 1 Aug 2022 · 0 repositories · arXiv:2208.00808
-
Mitigating Off-Policy Bias in Actor-Critic Methods with One-Step Q-learning: A Novel Correction Approach 1 Aug 2022 · 1 repository · arXiv:2208.00755
-
Playing a 2D Game Indefinitely using NEAT and Reinforcement Learning 28 Jul 2022 · 0 repositories · arXiv:2207.14140
-
Structural Similarity for Improved Transfer in Reinforcement Learning 27 Jul 2022 · 0 repositories · arXiv:2207.13813
-
Finite-Time Analysis of Asynchronous Q-learning under Diminishing Step-Size from Control-Theoretic View 25 Jul 2022 · 0 repositories · arXiv:2207.12217
-
Multi-Source AoI-Constrained Resource Minimization under HARQ: Heterogeneous Sampling Processes 19 Jul 2022 · 0 repositories · arXiv:2207.08996
-
On Decentralizing Federated Reinforcement Learning in Multi-Robot Scenarios 19 Jul 2022 · 0 repositories · arXiv:2207.09372
-
A Deep Reinforcement Learning Approach for Finding Non-Exploitable Strategies in Two-Player Atari Games 18 Jul 2022 · 2 repositories · arXiv:2207.08894Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 10 harvested samples)
-
Boolean Decision Rules for Reinforcement Learning Policy Summarisation 18 Jul 2022 · 0 repositories · arXiv:2207.08651
-
DDPG Learning for Aerial RIS-Assisted MU-MISO Communications 13 Jul 2022 · 0 repositories · arXiv:2207.06064
-
Multi-objective Optimization of Notifications Using Offline Reinforcement Learning 7 Jul 2022 · 0 repositories · arXiv:2207.03029
-
q-Learning in Continuous Time 2 Jul 2022 · 0 repositories · arXiv:2207.00713
-
Interactive Learning from Natural Language and Demonstrations using Signal Temporal Logic 1 Jul 2022 · 0 repositories · arXiv:2207.00627
-
Deep Reinforcement Learning with Swin Transformers 30 Jun 2022 · 1 repository · arXiv:2206.15269
-
On the Learning and Learnability of Quasimetrics 30 Jun 2022 · 2 repositories · arXiv:2206.15478Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples)
-
Recursive Reinforcement Learning 23 Jun 2022 · 0 repositories · arXiv:2206.11430
-
Reinforcement Learning under Partial Observability Guided by Learned Environment Models 23 Jun 2022 · 0 repositories · arXiv:2206.11708
-
Multi-Agent Car Parking using Reinforcement Learning 22 Jun 2022 · 1 repository · arXiv:2206.13338
-
Federated Stochastic Approximation under Markov Noise and Heterogeneity: Applications in Reinforcement Learning 21 Jun 2022 · 0 repositories · arXiv:2206.10185
-
The Integration of Machine Learning into Automated Test Generation: A Systematic Mapping Study 21 Jun 2022 · 0 repositories · arXiv:2206.10210
-
DNA: Proximal Policy Optimization with a Dual Network Architecture 20 Jun 2022 · 1 repository · arXiv:2206.10027
-
MASER: Multi-Agent Reinforcement Learning with Subgoals Generated from Experience Replay Buffer 20 Jun 2022 · 1 repository · arXiv:2206.10607
-
Sampling Efficient Deep Reinforcement Learning through Preference-Guided Stochastic Exploration 20 Jun 2022 · 1 repository · arXiv:2206.09627
-
Visual Radial Basis Q-Network 14 Jun 2022 · 0 repositories · arXiv:2206.06712
-
RL-GA: A Reinforcement Learning-Based Genetic Algorithm for Electromagnetic Detection Satellite Scheduling Problem 12 Jun 2022 · 0 repositories · arXiv:2206.05694
-
An Optimization Method-Assisted Ensemble Deep Reinforcement Learning Algorithm to Solve Unit Commitment Problems 9 Jun 2022 · 0 repositories · arXiv:2206.04249
-
Mildly Conservative Q-Learning for Offline Reinforcement Learning 9 Jun 2022 · 3 repositories · arXiv:2206.04745Syntology official: no sample here; runs from other or unrecorded repositories · 13 ran (of which 6 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 1 violated, 12 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 16 harvested samples) · 11 pointer-only (licence)
-
Concentration bounds for SSP Q-learning for average cost MDPs 7 Jun 2022 · 0 repositories · arXiv:2206.03328
-
DeepTPI: Test Point Insertion with Deep Reinforcement Learning 7 Jun 2022 · 1 repository · arXiv:2206.06975
-
Balancing Profit, Risk, and Sustainability for Portfolio Management 6 Jun 2022 · 0 repositories · arXiv:2207.02134
-
Goal-Space Planning with Subgoal Models 6 Jun 2022 · 0 repositories · arXiv:2206.02902
-
Offline RL for Natural Language Generation with Implicit Language Q Learning 5 Jun 2022 · 2 repositories · arXiv:2206.11871Syntology 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples) · 1 pointer-only (licence)
-
The Phenomenon of Policy Churn 1 Jun 2022 · 0 repositories · arXiv:2206.00730
-
Designing Rewards for Fast Learning 30 May 2022 · 0 repositories · arXiv:2205.15400
-
Efficient Reward Poisoning Attacks on Online Deep Reinforcement Learning 30 May 2022 · 1 repository · arXiv:2205.14842Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Deep Reinforcement Learning for Distributed and Uncoordinated Cognitive Radios Resource Allocation 27 May 2022 · 0 repositories · arXiv:2205.13944
-
Improving Bidding and Playing Strategies in the Trick-Taking game Wizard using Deep Q-Networks 27 May 2022 · 0 repositories · arXiv:2205.13834
-
An Experimental Comparison Between Temporal Difference and Residual Gradient with Neural Network Approximation 25 May 2022 · 0 repositories · arXiv:2205.12770
-
Deep Reinforcement Learning for Multi-class Imbalanced Training 24 May 2022 · 1 repository · arXiv:2205.12070
-
Multiple Domain Cyberspace Attack and Defense Game Based on Reward Randomization Reinforcement Learning 23 May 2022 · 0 repositories · arXiv:2205.10990
-
Optimizing Returns Using the Hurst Exponent and Q Learning on Momentum and Mean Reversion Strategies 23 May 2022 · 0 repositories · arXiv:2205.11122
-
Reinforced Pedestrian Attribute Recognition with Group Optimization Reward 21 May 2022 · 0 repositories · arXiv:2205.14042
-
Long Run Incremental Cost (LRIC) Distribution Network Pricing in UK, advising China's Distribution Network 20 May 2022 · 0 repositories · arXiv:2205.09946
-
Parallel bandit architecture based on laser chaos for reinforcement learning 19 May 2022 · 0 repositories · arXiv:2205.09543
-
Enforcing KL Regularization in General Tsallis Entropy Reinforcement Learning via Advantage Learning 16 May 2022 · 0 repositories · arXiv:2205.07885
-
Efficient Off-Policy Reinforcement Learning via Brain-Inspired Computing 14 May 2022 · 0 repositories · arXiv:2205.06978
-
Representation Learning for Context-Dependent Decision-Making 12 May 2022 · 0 repositories · arXiv:2205.05820
-
Characterizing the Action-Generalization Gap in Deep Q-Learning 11 May 2022 · 0 repositories · arXiv:2205.05588
-
Final Iteration Convergence Bound of Q-Learning: Switching System Approach 11 May 2022 · 0 repositories · arXiv:2205.05455
-
Neuromimetic Linear Systems -- Resilience and Learning 10 May 2022 · 0 repositories · arXiv:2205.05013
-
Simultaneous Double Q-learning with Conservative Advantage Learning for Actor-Critic Methods 8 May 2022 · 1 repository · arXiv:2205.03819
-
Chemoreception and chemotaxis of a three-sphere swimmer 5 May 2022 · 0 repositories · arXiv:2205.02678
-
Learning Value Functions from Undirected State-only Experience 26 Apr 2022 · 0 repositories · arXiv:2204.12458
-
EPiDA: An Easy Plug-in Data Augmentation Framework for High Performance Text Classification 24 Apr 2022 · 1 repository · arXiv:2204.11205Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 6 harvested samples) · 2 pointer-only (licence)
-
Graph Neural Network based Agent in Google Research Football 23 Apr 2022 · 0 repositories · arXiv:2204.11142
-
Learning how to Interact with a Complex Interface using Hierarchical Reinforcement Learning 21 Apr 2022 · 0 repositories · arXiv:2204.10374
-
Provably Efficient Kernelized Q-Learning 21 Apr 2022 · 0 repositories · arXiv:2204.10349
-
Joint Learning of Reward Machines and Policies in Environments with Partially Known Semantics 20 Apr 2022 · 0 repositories · arXiv:2204.11833
-
Reinforcement Re-ranking with 2D Grid-based Recommendation Panels 11 Apr 2022 · 0 repositories · arXiv:2204.04954
-
Optimizing the Long-Term Behaviour of Deep Reinforcement Learning for Pushing and Grasping 7 Apr 2022 · 0 repositories · arXiv:2204.03487
-
DouZero+: Improving DouDizhu AI by Opponent Modeling and Coach-guided Learning 6 Apr 2022 · 1 repository · arXiv:2204.02558
-
GAIL-PT: A Generic Intelligent Penetration Testing Framework with Generative Adversarial Imitation Learning 5 Apr 2022 · 1 repository · arXiv:2204.01975
-
REM: Routing Entropy Minimization for Capsule Networks 4 Apr 2022 · 0 repositories · arXiv:2204.01298
-
Deep Q-learning of global optimizer of multiply model parameters for viscoelastic imaging 1 Apr 2022 · 0 repositories · arXiv:2204.01844
-
Hysteresis-Based RL: Robustifying Reinforcement Learning-based Control Policies via Hybrid Control 1 Apr 2022 · 2 repositories · arXiv:2204.00654
-
Functional Stability of Discounted Markov Decision Processes Using Economic MPC Dissipativity Theory 31 Mar 2022 · 0 repositories · arXiv:2203.16989
-
Neural Q-learning for solving PDEs 31 Mar 2022 · 0 repositories · arXiv:2203.17128
-
Investigating the Properties of Neural Network Representations in Reinforcement Learning 30 Mar 2022 · 0 repositories · arXiv:2203.15955
-
Topological Experience Replay 29 Mar 2022 · 1 repository · arXiv:2203.15845
-
MERLIN -- Malware Evasion with Reinforcement LearnINg 24 Mar 2022 · 0 repositories · arXiv:2203.12980
-
The state-of-the-art review on resource allocation problem using artificial intelligence methods on various computing paradigms 23 Mar 2022 · 0 repositories · arXiv:2203.12315
-
A Note on Target Q-learning For Solving Finite MDPs with A Generative Oracle 22 Mar 2022 · 0 repositories · arXiv:2203.11489
-
Action Candidate Driven Clipped Double Q-learning for Discrete and Continuous Action Tasks 22 Mar 2022 · 1 repository · arXiv:2203.11526
-
Distributed Learning for Vehicular Dynamic Spectrum Access in Autonomous Driving 22 Mar 2022 · 0 repositories · arXiv:2204.10179
-
Does DQN really learn? Exploring adversarial training schemes in Pong 20 Mar 2022 · 0 repositories · arXiv:2203.10614
-
Infinite-Horizon Reach-Avoid Zero-Sum Games via Deep Reinforcement Learning 18 Mar 2022 · 0 repositories · arXiv:2203.10142
-
Orchestrated Value Mapping for Reinforcement Learning 14 Mar 2022 · 1 repository · arXiv:2203.07171
-
Reinforcement Learning for Optimal Control of a District Cooling Energy Plant 14 Mar 2022 · 0 repositories · arXiv:2203.07500
-
The Efficacy of Pessimism in Asynchronous Q-Learning 14 Mar 2022 · 0 repositories · arXiv:2203.07368
-
Random Ensemble Reinforcement Learning for Traffic Signal Control 10 Mar 2022 · 0 repositories · arXiv:2203.05961
-
Graph-based Reinforcement Learning meets Mixed Integer Programs: An application to 3D robot assembly discovery 8 Mar 2022 · 0 repositories · arXiv:2203.04120
-
Scalable multi-agent reinforcement learning for distributed control of residential energy flexibility 7 Mar 2022 · 0 repositories · arXiv:2203.03417
-
Deep Reinforcement Learning based Model-free On-line Dynamic Multi-Microgrid Formation to Enhance Resilience 6 Mar 2022 · 0 repositories · arXiv:2203.03030
-
Offline Deep Reinforcement Learning for Dynamic Pricing of Consumer Credit 6 Mar 2022 · 0 repositories · arXiv:2203.03003
-
A Learning Based Framework for Handling Uncertain Lead Times in Multi-Product Inventory Management 2 Mar 2022 · 0 repositories · arXiv:2203.00885
-
Improving the Diversity of Bootstrapped DQN by Replacing Priors With Noise 2 Mar 2022 · 0 repositories · arXiv:2203.01004
-
Pessimistic Q-Learning for Offline Reinforcement Learning: Towards Optimal Sample Complexity 28 Feb 2022 · 0 repositories · arXiv:2202.13890
-
Autonomous Warehouse Robot using Deep Q-Learning 21 Feb 2022 · 0 repositories · arXiv:2202.10019
-
PooL: Pheromone-inspired Communication Framework forLarge Scale Multi-Agent Reinforcement Learning 20 Feb 2022 · 0 repositories · arXiv:2202.09722
-
Retrieval-Augmented Reinforcement Learning 17 Feb 2022 · 0 repositories · arXiv:2202.08417
-
Goal Recognition as Reinforcement Learning 13 Feb 2022 · 1 repository · arXiv:2202.06356
-
Regularized Q-learning 11 Feb 2022 · 0 repositories · arXiv:2202.05404