Browse State-of-the-Art › Reinforcement Learning (RL) › Papers, page 62
Reinforcement Learning (RL)
Papers archive 2025-07-28
archive papers tagged: 15,113 · with a code link: 4,749 · where Syntology ran a sample: 1,416 (1,186 with a run with no instrument failure, 230 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (1,416 of 15,113 tagged: 1,186 with a run with no instrument failure, 230 where every run was a failure of Syntology's instrument)
Page 62 of 152: papers 6,101 to 6,200 of 15,113, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
By Fair Means or Foul: Quantifying Collusion in a Market Simulation with Deep Reinforcement Learning4 Jun 2024 0 repositories listed
-
FightLadder: A Benchmark for Competitive Multi-Agent Reinforcement Learning4 Jun 2024 0 repositories listed
-
iQRL -- Implicitly Quantized Representations for Sample-efficient Reinforcement Learning4 Jun 2024 0 repositories listed
-
Rectifying Reinforcement Learning for Reward Matching4 Jun 2024 0 repositories listed
-
Reinforcement Learning with Lookahead Information4 Jun 2024 0 repositories listed
-
Smaller Batches, Bigger Gains? Investigating the Impact of Batch Sizes on Reinforcement Learning Based Real-World Production Scheduling4 Jun 2024 0 repositories listed
-
A Fast Convergence Theory for Offline Decision Making3 Jun 2024 0 repositories listed
-
Causal prompting model-based offline reinforcement learning3 Jun 2024 0 repositories listed
-
Combinatorial Multivariant Multi-Armed Bandits with Applications to Episodic Reinforcement Learning and Beyond3 Jun 2024 0 repositories listed
-
Federated Learning-based Collaborative Wideband Spectrum Sensing and Scheduling for UAVs in UTM Systems3 Jun 2024 0 repositories listed
-
Learning the Target Network in Function Space3 Jun 2024 0 repositories listed
-
MOT: A Mixture of Actors Reinforcement Learning Method by Optimal Transport for Algorithmic Trading3 Jun 2024 0 repositories listed
-
NeoRL: Efficient Exploration for Nonepisodic RL3 Jun 2024 0 repositories listed
-
Reinforcement Learning as a Robotics-Inspired Framework for Insect Navigation: From Spatial Representations to Neural Implementation3 Jun 2024 0 repositories listed
-
REvolve: Reward Evolution with Large Language Models using Human Feedback3 Jun 2024 0 repositories listed
-
The Importance of Online Data: Understanding Preference Fine-tuning via Coverage3 Jun 2024 0 repositories listed
-
A Digital Twin Framework for Reinforcement Learning with Real-Time Self-Improvement via Human Assistive Teleoperation2 Jun 2024 0 repositories listed
-
Model Predictive Control and Reinforcement Learning: A Unified Framework Based on Dynamic Programming2 Jun 2024 0 repositories listed
-
Decision Mamba: Reinforcement Learning via Hybrid Selective Sequence Modeling31 May 2024 0 repositories listed
-
Reinforcement Learning for Sociohydrology31 May 2024 0 repositories listed
-
Bilevel reinforcement learning via the development of hyper-gradient without lower-level convexity30 May 2024 0 repositories listed
-
Efficient Stimuli Generation using Reinforcement Learning in Design Verification30 May 2024 0 repositories listed
-
From Words to Actions: Unveiling the Theoretical Underpinnings of LLM-Driven Autonomous Systems30 May 2024 0 repositories listed
-
Hybrid Reinforcement Learning Framework for Mixed-Variable Problems30 May 2024 0 repositories listed
-
Learning to Discuss Strategically: A Case Study on One Night Ultimate Werewolf30 May 2024 0 repositories listed
-
SleeperNets: Universal Backdoor Poisoning Attacks Against Reinforcement Learning Agents30 May 2024 0 repositories listed
-
Policy Zooming: Adaptive Discretization-based Infinite-Horizon Average-Reward Reinforcement Learning29 May 2024 0 repositories listed
-
Preferred-Action-Optimized Diffusion Policies for Offline Reinforcement Learning29 May 2024 0 repositories listed
-
Safety through Permissibility: Shield Construction for Fast and Safe Reinforcement Learning29 May 2024 0 repositories listed
-
Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF29 May 2024 0 repositories listed
-
Extreme Value Monte Carlo Tree Search28 May 2024 0 repositories listed
-
Highway Reinforcement Learning28 May 2024 0 repositories listed
-
Mollification Effects of Policy Gradient Methods28 May 2024 0 repositories listed
-
Rethinking Pruning for Backdoor Mitigation: An Optimization Perspective28 May 2024 0 repositories listed
-
LeDex: Training LLMs to Better Self-Debug and Explain Code28 May 2024 0 repositories listed
-
Biological Neurons Compete with Deep Reinforcement Learning in Sample Efficiency in a Simulated Gameworld27 May 2024 0 repositories listed
-
Ontology-Enhanced Decision-Making for Autonomous Agents in Dynamic and Partially Observable Environments27 May 2024 0 repositories listed
-
Oracle-Efficient Reinforcement Learning for Max Value Ensembles27 May 2024 0 repositories listed
-
Structured Graph Network for Constrained Robot Crowd Navigation with Low Fidelity Simulation27 May 2024 0 repositories listed
-
Trajectory Data Suffices for Statistically Efficient Learning in Offline RL with Linear q^π-Realizability and Concentrability27 May 2024 0 repositories listed
-
An Evolutionary Framework for Connect-4 as Test-Bed for Comparison of Advanced Minimax, Q-Learning and MCTS26 May 2024 0 repositories listed
-
Fast TRAC: A Parameter-Free Optimizer for Lifelong Reinforcement Learning26 May 2024 0 repositories listed
-
Reinforcement Learning for Jump-Diffusions, with Financial Applications26 May 2024 0 repositories listed
-
AIGB: Generative Auto-bidding via Conditional Diffusion Modeling25 May 2024 0 repositories listed
-
25 May 2024 0 repositories listed Syntology 7 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples) · 1 pointer-only (licence)
-
Cooperative Backdoor Attack in Decentralized Reinforcement Learning with Theoretical Guarantee24 May 2024 0 repositories listed
-
Extracting Heuristics from Large Language Models for Reward Shaping in Reinforcement Learning24 May 2024 0 repositories listed
-
Embedding-Aligned Language Models24 May 2024 0 repositories listed
-
Human-in-the-loop Reinforcement Learning for Data Quality Monitoring in Particle Physics Experiments24 May 2024 0 repositories listed
-
Knowledge-Informed Auto-Penetration Testing Based on Reinforcement Learning with Reward Machine24 May 2024 0 repositories listed
-
SF-DQN: Provable Knowledge Transfer using Successor Feature for Deep Reinforcement Learning24 May 2024 0 repositories listed
-
TrojanForge: Generating Adversarial Hardware Trojan Examples Using Reinforcement Learning24 May 2024 0 repositories listed
-
A finite time analysis of distributed Q-learning23 May 2024 0 repositories listed
-
Blood Glucose Control Via Pre-trained Counterfactual Invertible Neural Networks23 May 2024 0 repositories listed
-
Exclusively Penalized Q-learning for Offline Reinforcement Learning23 May 2024 0 repositories listed
-
Policy Gradient Methods for Risk-Sensitive Distributional Reinforcement Learning with Provable Convergence23 May 2024 0 repositories listed
-
Efficiently Training Deep-Learning Parametric Policies using Lagrangian Duality23 May 2024 0 repositories listed
-
Autonomous Algorithm for Training Autonomous Vehicles with Minimal Human Intervention22 May 2024 0 repositories listed
-
HighwayLLM: Decision-Making and Navigation in Highway Driving with RL-Informed Language Model22 May 2024 0 repositories listed
-
Large Language Models (LLMs) Assisted Wireless Network Deployment in Urban Settings22 May 2024 0 repositories listed
-
Leader Reward for POMO-Based Neural Combinatorial Optimization22 May 2024 0 repositories listed
-
A Multimodal Learning-based Approach for Autonomous Landing of UAV21 May 2024 0 repositories listed
-
Multi-Agent Reinforcement Learning with Hierarchical Coordination for Emergency Responder Stationing21 May 2024 0 repositories listed
-
Practical and efficient quantum circuit synthesis and transpiling with Reinforcement Learning21 May 2024 0 repositories listed
-
Rethinking Robustness Assessment: Adversarial Attacks on Learning-based Quadrupedal Locomotion Controllers21 May 2024 0 repositories listed
-
Investigating the Impact of Choice on Deep Reinforcement Learning for Space Controls20 May 2024 0 repositories listed
-
Learning Future Representation with Synthetic Observations for Sample-efficient Reinforcement Learning20 May 2024 0 repositories listed
-
Comparisons Are All You Need for Optimizing Smooth Functions19 May 2024 0 repositories listed
-
Do No Harm: A Counterfactual Approach to Safe Reinforcement Learning19 May 2024 0 repositories listed
-
Combined film and pulse heating of lithium ion batteries to improve performance in low ambient temperature18 May 2024 0 repositories listed
-
Optimal control barrier functions for RL based safe powertrain control18 May 2024 0 repositories listed
-
Towards Robust Policy: Enhancing Offline Reinforcement Learning with Adversarial Attacks and Defenses18 May 2024 0 repositories listed
-
LLM-based Multi-Agent Reinforcement Learning: Current and Future Directions17 May 2024 0 repositories listed
-
Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement Learning16 May 2024 0 repositories listed
-
Stochastic Q-learning for Large Discrete Action Spaces16 May 2024 0 repositories listed
-
Deep Learning in Earthquake Engineering: A Comprehensive Review15 May 2024 0 repositories listed
-
Fast Two-Time-Scale Stochastic Gradient Method with Applications in Reinforcement Learning15 May 2024 0 repositories listed
-
IM-RAG: Multi-Round Retrieval-Augmented Generation Through Learning Inner Monologues15 May 2024 0 repositories listed
-
Deep Reinforcement Learning for Real-Time Ground Delay Program Revision and Corresponding Flight Delay Assignments14 May 2024 0 repositories listed
-
vMFER: Von Mises-Fisher Experience Resampling Based on Uncertainty of Gradient Directions for Policy Improvement14 May 2024 0 repositories listed
-
Intrinsic Rewards for Exploration without Harm from Observational Noise: A Simulation Study Based on the Free Energy Principle13 May 2024 0 repositories listed
-
Near-Optimal Regret in Linear MDPs with Aggregate Bandit Feedback13 May 2024 0 repositories listed
-
Neural Network Compression for Reinforcement Learning Tasks13 May 2024 0 repositories listed
-
Reducing Risk for Assistive Reinforcement Learning Policies with Diffusion Models13 May 2024 0 repositories listed
-
Ensemble Successor Representations for Task Generalization in Offline-to-Online Reinforcement Learning12 May 2024 0 repositories listed
-
Fairness in Reinforcement Learning: A Survey11 May 2024 0 repositories listed
-
Dominion: A New Frontier for AI Research10 May 2024 0 repositories listed
-
Improving Targeted Molecule Generation through Language Model Fine-Tuning Via Reinforcement Learning10 May 2024 0 repositories listed
-
Space Processor Computation Time Analysis for Reinforcement Learning and Run Time Assurance Control Policies10 May 2024 0 repositories listed
-
An Overview of Machine Learning-Enabled Optimization for Reconfigurable Intelligent Surfaces-Aided 6G Networks: From Reinforcement Learning to Large Language Models9 May 2024 0 repositories listed
-
Fast Stochastic Policy Gradient: Negative Momentum for Reinforcement Learning8 May 2024 0 repositories listed
-
Genetic Drift Regularization: on preventing Actor Injection from breaking Evolution Strategies7 May 2024 0 repositories listed
-
Improving Offline Reinforcement Learning with Inaccurate Simulators7 May 2024 0 repositories listed
-
Roadside Units Assisted Localized Automated Vehicle Maneuvering: An Offline Reinforcement Learning Approach7 May 2024 0 repositories listed
-
Out-of-Distribution Adaptation in Offline RL: Counterfactual Reasoning via Causal Normalizing Flows6 May 2024 0 repositories listed
-
Safe Reinforcement Learning with Learned Non-Markovian Safety Constraints5 May 2024 0 repositories listed
-
UDUC: An Uncertainty-driven Approach for Learning-based Robust Control4 May 2024 0 repositories listed
-
A Model-based Multi-Agent Personalized Short-Video Recommender System3 May 2024 0 repositories listed
-
Learning Optimal Deterministic Policies with Stochastic Policy Gradients3 May 2024 0 repositories listed
Syntology lines on 2 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.