Browse State-of-the-Art › Reinforcement Learning › Papers, page 57
Reinforcement Learning
Papers archive 2025-07-28
archive papers tagged: 13,178 · with a code link: 4,183 · where Syntology ran a sample: 1,175 (988 with a run with no instrument failure, 187 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (1,175 of 13,178 tagged: 988 with a run with no instrument failure, 187 where every run was a failure of Syntology's instrument)
Page 57 of 132: papers 5,601 to 5,700 of 13,178, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Proxy-RLHF: Decoupling Generation and Alignment in Large Language Model with Proxy7 Mar 2024 0 repositories listed
-
Teaching Large Language Models to Reason with Reinforcement Learning7 Mar 2024 0 repositories listed
-
Why Online Reinforcement Learning is Causal7 Mar 2024 0 repositories listed
-
A Survey on Applications of Reinforcement Learning in Spatial Resource Allocation6 Mar 2024 0 repositories listed
-
Reconciling Reality through Simulation: A Real-to-Sim-to-Real Approach for Robust Manipulation6 Mar 2024 0 repositories listed
-
A Zero-Shot Reinforcement Learning Strategy for Autonomous Guidewire Navigation5 Mar 2024 0 repositories listed
-
Autonomous vehicle decision and control through reinforcement learning with traffic flow randomization5 Mar 2024 0 repositories listed
-
Enhancing LLM Safety via Constrained Direct Preference Optimization4 Mar 2024 0 repositories listed
-
Koopman-Assisted Reinforcement Learning4 Mar 2024 0 repositories listed
-
Tsallis Entropy Regularization for Linearly Solvable MDP and Linear Quadratic Regulator4 Mar 2024 0 repositories listed
-
Twisting Lids Off with Two Hands4 Mar 2024 0 repositories listed
-
Sample Efficient Myopic Exploration Through Multitask Reinforcement Learning with Diverse Tasks3 Mar 2024 0 repositories listed
-
Towards Provable Log Density Policy Gradient3 Mar 2024 0 repositories listed
-
Continuous Mean-Zero Disagreement-Regularized Imitation Learning (CMZ-DRIL)2 Mar 2024 0 repositories listed
-
On the Role of Information Structure in Reinforcement Learning for Partially-Observable Sequential Teams and Games1 Mar 2024 0 repositories listed
-
Scale-free Adversarial Reinforcement Learning1 Mar 2024 0 repositories listed
-
Conflict-Averse Gradient Aggregation for Constrained Multi-Objective Reinforcement Learning1 Mar 2024 0 repositories listed
-
SELFI: Autonomous Self-Improvement with Reinforcement Learning for Social Navigation1 Mar 2024 0 repositories listed
-
Adaptive Testing Environment Generation for Connected and Automated Vehicles with Dense Reinforcement Learning29 Feb 2024 0 repositories listed
-
Deep Reinforcement Learning: A Convex Optimization Approach29 Feb 2024 0 repositories listed
-
Disentangling the Causes of Plasticity Loss in Neural Networks29 Feb 2024 0 repositories listed
-
RL-GPT: Integrating Reinforcement Learning and Code-as-policy29 Feb 2024 0 repositories listed
-
Human-Centric Aware UAV Trajectory Planning in Search and Rescue Missions Employing Multi-Objective Reinforcement Learning with AHP and Similarity-Based Experience Replay28 Feb 2024 0 repositories listed
-
Implementing Online Reinforcement Learning with Clustering Neural Networks28 Feb 2024 0 repositories listed
-
Provable Risk-Sensitive Distributional Reinforcement Learning with General Function Approximation28 Feb 2024 0 repositories listed
-
Provably Efficient Partially Observable Risk-Sensitive Reinforcement Learning with Hindsight Observation28 Feb 2024 0 repositories listed
-
Why Do Animals Need Shaping? A Theory of Task Composition and Curriculum Learning28 Feb 2024 0 repositories listed
-
Advancing Investment Frontiers: Industry-grade Deep Reinforcement Learning for Portfolio Optimization27 Feb 2024 0 repositories listed
-
Multi-Agent Deep Reinforcement Learning for Distributed Satellite Routing27 Feb 2024 0 repositories listed
-
Enhancing Kubernetes Automated Scheduling with Deep Learning and Reinforcement Techniques for Large-Scale Cloud Computing Optimization26 Feb 2024 0 repositories listed
-
Monitoring Fidelity of Online Reinforcement Learning Algorithms in Clinical Trials26 Feb 2024 0 repositories listed
-
Program-Based Strategy Induction for Reinforcement Learning26 Feb 2024 0 repositories listed
-
QF-tuner: Breaking Tradition in Reinforcement Learning26 Feb 2024 0 repositories listed
-
Reinforcement Learning Jazz Improvisation: When Music Meets Game Theory25 Feb 2024 0 repositories listed
-
Analysis of Off-Policy Multi-Step TD-Learning with Linear Function Approximation24 Feb 2024 0 repositories listed
-
Safety Optimized Reinforcement Learning via Multi-Objective Policy Optimization23 Feb 2024 0 repositories listed
-
Shapley Value Based Multi-Agent Reinforcement Learning: Theory, Method and Its Application to Energy Network23 Feb 2024 0 repositories listed
-
Trajectory-wise Iterative Reinforcement Learning Framework for Auto-bidding23 Feb 2024 0 repositories listed
-
Applying Reinforcement Learning to Optimize Traffic Light Cycles22 Feb 2024 0 repositories listed
-
Model-Based Reinforcement Learning Control of Reaction-Diffusion Problems22 Feb 2024 0 repositories listed
-
MR-ARL: Model Reference Adaptive Reinforcement Learning for Robustly Stable On-Policy Data-Driven LQR22 Feb 2024 0 repositories listed
-
Reinforcement Learning with Elastic Time Steps22 Feb 2024 0 repositories listed
-
AttackGNN: Red-Teaming GNNs in Hardware Security Using Reinforcement Learning21 Feb 2024 0 repositories listed
-
Reinforcement learning-assisted quantum architecture search for variational quantum algorithms21 Feb 2024 0 repositories listed
-
Synthesis of Hierarchical Controllers Based on Deep Reinforcement Learning Policies21 Feb 2024 0 repositories listed
-
Uniform Last-Iterate Guarantee for Bandits and Reinforcement Learning20 Feb 2024 0 repositories listed
-
Advancing Monocular Video-Based Gait Analysis Using Motion Imitation with Physics-Based Simulation20 Feb 2024 0 repositories listed
-
Deep Hedging with Market Impact20 Feb 2024 0 repositories listed
-
Evolutionary Reinforcement Learning: A Systematic Review and Future Directions20 Feb 2024 0 repositories listed
-
CovRL: Fuzzing JavaScript Engines with Coverage-Guided Reinforcement Learning for LLM-based Mutation19 Feb 2024 0 repositories listed
-
In value-based deep reinforcement learning, a pruned network is a good network19 Feb 2024 0 repositories listed
-
Reinforcement Learning for Optimal Execution when Liquidity is Time-Varying19 Feb 2024 0 repositories listed
-
Interactive Garment Recommendation with User in the Loop18 Feb 2024 0 repositories listed
-
Programmatic Reinforcement Learning: Navigating Gridworlds18 Feb 2024 0 repositories listed
-
Implementation of a Model of the Cortex Basal Ganglia Loop17 Feb 2024 0 repositories listed
-
Multi Task Inverse Reinforcement Learning for Common Sense Reward17 Feb 2024 0 repositories listed
-
Reinforcement learning to maximise wind turbine energy generation17 Feb 2024 0 repositories listed
-
Large Scale Constrained Clustering With Reinforcement Learning15 Feb 2024 0 repositories listed
-
Reinforcement Learning for Solving Stochastic Vehicle Routing Problem with Time Windows15 Feb 2024 0 repositories listed
-
Universal Black-Box Reward Poisoning Attack against Offline Reinforcement Learning15 Feb 2024 0 repositories listed
-
Smart Information Exchange for Unsupervised Federated Learning via Reinforcement Learning15 Feb 2024 0 repositories listed
-
Learning Interpretable Policies in Hindsight-Observable POMDPs through Partially Supervised Reinforcement Learning14 Feb 2024 0 repositories listed
-
LL-GABR: Energy Efficient Live Video Streaming Using Reinforcement Learning14 Feb 2024 0 repositories listed
-
How does Your RL Agent Explore? An Optimal Transport Analysis of Occupancy Measure Trajectories14 Feb 2024 0 repositories listed
-
PMGDA: A Preference-based Multiple Gradient Descent Algorithm14 Feb 2024 0 repositories listed
-
Reinforcement Learning from Human Feedback with Active Queries14 Feb 2024 0 repositories listed
-
Steady-State Error Compensation for Reinforcement Learning with Quadratic Rewards14 Feb 2024 0 repositories listed
-
Towards Robust Model-Based Reinforcement Learning Against Adversarial Corruption14 Feb 2024 0 repositories listed
-
Uncertainty-Aware Transient Stability-Constrained Preventive Redispatch: A Distributional Reinforcement Learning Approach14 Feb 2024 0 repositories listed
-
IR-Aware ECO Timing Optimization Using Reinforcement Learning12 Feb 2024 0 repositories listed
-
Large Language Models as Agents in Two-Player Games12 Feb 2024 0 repositories listed
-
MAIDCRL: Semi-centralized Multi-Agent Influence Dense-CNN Reinforcement Learning12 Feb 2024 0 repositories listed
-
Measurement Scheduling for ICU Patients with Offline Reinforcement Learning12 Feb 2024 0 repositories listed
-
Near-Minimax-Optimal Distributional Reinforcement Learning with a Generative Model12 Feb 2024 0 repositories listed
-
Future Prediction Can be a Strong Evidence of Good History Representation in Partially Observable Environments11 Feb 2024 0 repositories listed
-
Natural Language Reinforcement Learning11 Feb 2024 0 repositories listed
-
Towards Generalized Inverse Reinforcement Learning11 Feb 2024 0 repositories listed
-
Using Large Language Models to Automate and Expedite Reinforcement Learning with Reward Machine11 Feb 2024 0 repositories listed
-
Principled Penalty-based Methods for Bilevel Reinforcement Learning and RLHF10 Feb 2024 0 repositories listed
-
Corruption Robust Offline Reinforcement Learning with Human Feedback9 Feb 2024 0 repositories listed
-
Hierarchical Transformers are Efficient Meta-Reinforcement Learners9 Feb 2024 0 repositories listed
-
High-Precision Geosteering via Reinforcement Learning and Particle Filters9 Feb 2024 0 repositories listed
-
Value function interference and greedy action selection in value-based multi-objective reinforcement learning9 Feb 2024 0 repositories listed
-
Differentially Private Deep Model-Based Reinforcement Learning8 Feb 2024 0 repositories listed
-
Attention-Enhanced Prioritized Proximal Policy Optimization for Adaptive Edge Caching8 Feb 2024 0 repositories listed
-
Federated Offline Reinforcement Learning: Collaborative Single-Policy Coverage Suffices8 Feb 2024 0 repositories listed
-
Offline Actor-Critic Reinforcement Learning Scales to Large Models8 Feb 2024 0 repositories listed
-
Scaling Artificial Intelligence for Digital Wargaming in Support of Decision-Making8 Feb 2024 0 repositories listed
-
A computational approach to visual ecology with deep reinforcement learning7 Feb 2024 0 repositories listed
-
Analyzing Adversarial Inputs in Deep Reinforcement Learning7 Feb 2024 0 repositories listed
-
Code as Reward: Empowering Reinforcement Learning with VLMs7 Feb 2024 0 repositories listed
-
6 Feb 2024 0 repositories listed Syntology 5 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 8 harvested samples) · 6 pointer-only (licence)
-
No-Regret Reinforcement Learning in Smooth MDPs6 Feb 2024 0 repositories listed
-
Reinforcement Learning from Bagged Reward6 Feb 2024 0 repositories listed
-
Transductive Reward Inference on Graph6 Feb 2024 0 repositories listed
-
A Multi-step Loss Function for Robust Learning of the Dynamics in Model-based Reinforcement Learning5 Feb 2024 0 repositories listed
-
A Reinforcement Learning Approach for Dynamic Rebalancing in Bike-Sharing System5 Feb 2024 0 repositories listed
-
Abstracted Trajectory Visualization for Explainability in Reinforcement Learning5 Feb 2024 0 repositories listed
-
Assessing the Impact of Distribution Shift on Reinforcement Learning Performance5 Feb 2024 0 repositories listed
-
Curriculum reinforcement learning for quantum architecture search under hardware errors5 Feb 2024 0 repositories listed
Syntology lines on 1 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.