Browse State-of-the-Art › Reinforcement Learning (RL) › Papers, page 55
Reinforcement Learning (RL)
Papers archive 2025-07-28
archive papers tagged: 15,113 · with a code link: 4,749 · where Syntology ran a sample: 1,416 (1,186 with a run with no instrument failure, 230 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (1,416 of 15,113 tagged: 1,186 with a run with no instrument failure, 230 where every run was a failure of Syntology's instrument)
Page 55 of 152: papers 5,401 to 5,500 of 15,113, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
RAPID: Robust and Agile Planner Using Inverse Reinforcement Learning for Vision-Based Drone Navigation4 Feb 2025 0 repositories listed
-
VolleyBots: A Testbed for Multi-Drone Volleyball Game Combining Motion Control and Strategic Play4 Feb 2025 0 repositories listed
-
ACECODER: Acing Coder RL via Automated Test-Case Synthesis3 Feb 2025 0 repositories listed
-
Dynamic object goal pushing with mobile manipulators through model-free constrained reinforcement learning3 Feb 2025 0 repositories listed
-
Preference VLM: Leveraging VLMs for Scalable Preference-Based Reinforcement Learning3 Feb 2025 0 repositories listed
-
Reinforcement Learning for Long-Horizon Interactive LLM Agents3 Feb 2025 0 repositories listed
-
Reinforcement Learning with Segment Feedback3 Feb 2025 0 repositories listed
-
Resilient UAV Trajectory Planning via Few-Shot Meta-Offline Reinforcement Learning3 Feb 2025 0 repositories listed
-
The Differences Between Direct Alignment Algorithms are a Blur3 Feb 2025 0 repositories listed
-
Toward Task Generalization via Memory Augmentation in Meta-Reinforcement Learning3 Feb 2025 0 repositories listed
-
Zeroth-order Informed Fine-Tuning for Diffusion Model: A Recursive Likelihood Ratio Optimizer2 Feb 2025 0 repositories listed
-
A Differentiated Reward Method for Reinforcement Learning based Multi-Vehicle Cooperative Decision-Making Algorithms1 Feb 2025 0 repositories listed
-
Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network1 Feb 2025 0 repositories listed
-
Model-Free Predictive Control: Introductory Algebraic Calculations, and a Comparison with HEOL and ANNs1 Feb 2025 0 repositories listed
-
Decorrelated Soft Actor-Critic for Efficient Deep Reinforcement Learning31 Jan 2025 0 repositories listed
-
O-MAPL: Offline Multi-agent Preference Learning31 Jan 2025 0 repositories listed
-
Optimizing Job Allocation using Reinforcement Learning with Graph Neural Networks31 Jan 2025 0 repositories listed
-
RLS3: RL-Based Synthetic Sample Selection to Enhance Spatial Reasoning in Vision-Language Models for Indoor Autonomous Perception31 Jan 2025 0 repositories listed
-
Towards Physiologically Sensible Predictions via the Rule-based Reinforcement Learning Layer31 Jan 2025 0 repositories listed
-
B3C: A Minimalist Approach to Offline Multi-Agent Reinforcement Learning30 Jan 2025 0 repositories listed
-
Model-Free RL Agents Demonstrate System 1-Like Intentionality30 Jan 2025 0 repositories listed
-
A Dual-Agent Adversarial Framework for Robust Generalization in Deep Reinforcement Learning29 Jan 2025 0 repositories listed
-
From Sparse to Dense: Toddler-inspired Reward Transition in Goal-Oriented Reinforcement Learning29 Jan 2025 0 repositories listed
-
Reinforcement-Learning Portfolio Allocation with Dynamic Embedding of Market Information29 Jan 2025 0 repositories listed
-
RL-based Query Rewriting with Distilled LLM for online E-Commerce Systems29 Jan 2025 0 repositories listed
-
Challenges in Ensuring AI Safety in DeepSeek-R1 Models: The Shortcomings of Reinforcement Learning Strategies28 Jan 2025 0 repositories listed
-
Exploratory Mean-Variance Portfolio Optimization with Regime-Switching Market Dynamics28 Jan 2025 0 repositories listed
-
Heterogeneity-aware Personalized Federated Learning via Adaptive Dual-Agent Reinforcement Learning28 Jan 2025 0 repositories listed
-
Improving Vision-Language-Action Model with Online Reinforcement Learning28 Jan 2025 0 repositories listed
-
Integrating Reinforcement Learning and AI Agents for Adaptive Robotic Interaction and Assistance in Dementia Care28 Jan 2025 0 repositories listed
-
Safe Reinforcement Learning for Real-World Engine Control28 Jan 2025 0 repositories listed
-
SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training28 Jan 2025 0 repositories listed
-
Flexible Blood Glucose Control: Offline Reinforcement Learning from Human Feedback27 Jan 2025 0 repositories listed
-
FuzzyLight: A Robust Two-Stage Fuzzy Approach for Traffic Signal Control Works in Real Cities27 Jan 2025 0 repositories listed
-
MPC4RL -- A Software Package for Reinforcement Learning based on Model Predictive Control27 Jan 2025 0 repositories listed
-
Selective Experience Sharing in Reinforcement Learning Enhances Interference Management27 Jan 2025 0 repositories listed
-
27 Jan 2025 0 repositories listed Syntology 5 ran (of which 4 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified (of 9 harvested samples) · 4 pointer-only (licence)
-
Learning-Enhanced Safeguard Control for High-Relative-Degree Systems: Robust Optimization under Disturbances and Faults26 Jan 2025 0 repositories listed
-
Data Center Cooling System Optimization Using Offline Reinforcement Learning25 Jan 2025 0 repositories listed
-
Faster Machine Translation Ensembling with Reinforcement Learning and Competitive Correction25 Jan 2025 0 repositories listed
-
Age and Power Minimization via Meta-Deep Reinforcement Learning in UAV Networks24 Jan 2025 0 repositories listed
-
Coordinating Ride-Pooling with Public Transit using Reward-Guided Conservative Q-Learning: An Offline Training and Online Fine-Tuning Reinforcement Learning Framework24 Jan 2025 0 repositories listed
-
Towards Efficient Multi-Objective Optimisation for Real-World Power Grid Topology Control24 Jan 2025 0 repositories listed
-
Large Language Model driven Policy Exploration for Recommender Systems23 Jan 2025 0 repositories listed
-
AdaWM: Adaptive World Model based Planning for Autonomous Driving22 Jan 2025 0 repositories listed
-
Deep Reinforcement Learning with Hybrid Intrinsic Reward Model22 Jan 2025 0 repositories listed
-
Evolution and The Knightian Blindspot of Machine Learning22 Jan 2025 0 repositories listed
-
Exploring the Technology Landscape through Topic Modeling, Expert Involvement, and Reinforcement Learning22 Jan 2025 0 repositories listed
-
MONA: Myopic Optimization with Non-myopic Approval Can Mitigate Multi-step Reward Hacking22 Jan 2025 0 repositories listed
-
On Generalization and Distributional Update for Mimicking Observations with Adequate Exploration22 Jan 2025 0 repositories listed
-
Reinforcement learning Based Automated Design of Differential Evolution Algorithm for Black-box Optimization22 Jan 2025 0 repositories listed
-
State Combinatorial Generalization In Decision Making With Conditional Diffusion Models22 Jan 2025 0 repositories listed
-
Extend Adversarial Policy Against Neural Machine Translation via Unknown Token21 Jan 2025 0 repositories listed
-
Reinforcement Learning Constrained Beam Search for Parameter Optimization of Paper Drying Under Flexible Constraints21 Jan 2025 0 repositories listed
-
RL-RC-DoT: A Block-level RL agent for Task-Aware Video Compression21 Jan 2025 0 repositories listed
-
RedStar: Does Scaling Long-CoT Data Unlock Better Slow-Reasoning Systems?20 Jan 2025 0 repositories listed
-
Solving Finite-Horizon MDPs via Low-Rank Tensors17 Jan 2025 0 repositories listed
-
From Explainability to Interpretability: Interpretable Policies in Reinforcement Learning Via Model Explanation16 Jan 2025 0 repositories listed
-
RE-POSE: Synergizing Reinforcement Learning-Based Partitioning and Offloading for Edge Object Detection16 Jan 2025 0 repositories listed
-
Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models16 Jan 2025 0 repositories listed
-
Average-Reward Reinforcement Learning with Entropy Regularization15 Jan 2025 0 repositories listed
-
Projection Implicit Q-Learning with Support Constraint for Offline Reinforcement Learning15 Jan 2025 0 repositories listed
-
Reinforcement Learning-Enhanced Procedural Generation for Dynamic Narrative-Driven AR Experiences15 Jan 2025 0 repositories listed
-
Decision Transformers for RIS-Assisted Systems with Diffusion Model-Based Channel Acquisition14 Jan 2025 0 repositories listed
-
FDPP: Fine-tune Diffusion Policy with Human Preference14 Jan 2025 0 repositories listed
-
Hybrid Action Based Reinforcement Learning for Multi-Objective Compatible Autonomous Driving14 Jan 2025 0 repositories listed
-
Combining LLM decision and RL action selection to improve RL policy for adaptive interventions13 Jan 2025 0 repositories listed
-
Future-Conditioned Recommendations with Multi-Objective Controllable Decision Transformer13 Jan 2025 0 repositories listed
-
RbRL2.0: Integrated Reward and Policy Learning for Rating-based Reinforcement Learning13 Jan 2025 0 repositories listed
-
Average Reward Reinforcement Learning for Wireless Radio Resource Management12 Jan 2025 0 repositories listed
-
DRDT3: Diffusion-Refined Decision Test-Time Training Model12 Jan 2025 0 repositories listed
-
Pareto Set Learning for Multi-Objective Reinforcement Learning12 Jan 2025 0 repositories listed
-
AlgoPilot: Fully Autonomous Program Synthesis Without Human-Written Programs11 Jan 2025 0 repositories listed
-
Hierarchical Reinforcement Learning for Optimal Agent Grouping in Cooperative Systems11 Jan 2025 0 repositories listed
-
Counterfactually Fair Reinforcement Learning via Sequential Data Preprocessing10 Jan 2025 0 repositories listed
-
Diffusion Models for Smarter UAVs: Decision-Making and Modeling10 Jan 2025 0 repositories listed
-
Investigating the Impact of Observation Space Design Choices On Training Reinforcement Learning Solutions for Spacecraft Problems10 Jan 2025 0 repositories listed
-
Real-Time Integrated Dispatching and Idle Fleet Steering with Deep Reinforcement Learning for A Meal Delivery Platform10 Jan 2025 0 repositories listed
-
LearningFlow: Automated Policy Learning Workflow for Urban Driving with Large Language Models9 Jan 2025 0 repositories listed
-
Deep Transfer Q-Learning for Offline Non-Stationary Reinforcement Learning8 Jan 2025 0 repositories listed
-
Risk-averse policies for natural gas futures trading using distributional reinforcement learning8 Jan 2025 0 repositories listed
-
Safe Reinforcement Learning with Minimal Supervision8 Jan 2025 0 repositories listed
-
Explainable Reinforcement Learning via Temporal Policy Decomposition7 Jan 2025 0 repositories listed
-
Run-and-tumble chemotaxis using reinforcement learning7 Jan 2025 0 repositories listed
-
Interpretable Recognition of Fused Magnesium Furnace Working Conditions with Deep Convolutional Stochastic Configuration Networks6 Jan 2025 0 repositories listed
-
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes6 Jan 2025 0 repositories listed
-
A New Interpretation of the Certainty-Equivalence Approach for PAC Reinforcement Learning with a Generative Model5 Jan 2025 0 repositories listed
-
AMM: Adaptive Modularized Reinforcement Model for Multi-city Traffic Signal Control5 Jan 2025 0 repositories listed
-
SR-Reward: Taking The Path More Traveled4 Jan 2025 0 repositories listed
-
On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures3 Jan 2025 0 repositories listed
-
Proposing Hierarchical Goal-Conditioned Policy Planning in Multi-Goal Reinforcement Learning3 Jan 2025 0 repositories listed
-
A Graphical Approach to State Variable Selection in Off-policy Learning1 Jan 2025 0 repositories listed
-
Neural Motion Simulator Pushing the Limit of World Models in Reinforcement Learning1 Jan 2025 0 repositories listed
-
RaSS: Improving Denoising Diffusion Samplers with Reinforced Active Sampling Scheduler1 Jan 2025 0 repositories listed
-
FORM: Learning Expressive and Transferable First-Order Logic Reward Machines31 Dec 2024 0 repositories listed
-
Towards Unraveling and Improving Generalization in World Models31 Dec 2024 0 repositories listed
-
An Unsupervised Anomaly Detection in Electricity Consumption Using Reinforcement Learning and Time Series Forest Based Framework30 Dec 2024 0 repositories listed
-
Isoperimetry is All We Need: Langevin Posterior Sampling for RL with Sublinear Regret30 Dec 2024 0 repositories listed
-
UnrealZoo: Enriching Photo-realistic Virtual Worlds for Embodied AI30 Dec 2024 0 repositories listed
-
Weber-Fechner Law in Temporal Difference learning derived from Control as Inference30 Dec 2024 0 repositories listed
Syntology lines on 1 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.