Browse State-of-the-Art › Reinforcement Learning (RL) › Papers, page 54
Reinforcement Learning (RL)
Papers archive 2025-07-28
archive papers tagged: 15,113 · with a code link: 4,749 · where Syntology ran a sample: 1,416 (1,186 with a run with no instrument failure, 230 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (1,416 of 15,113 tagged: 1,186 with a run with no instrument failure, 230 where every run was a failure of Syntology's instrument)
Page 54 of 152: papers 5,301 to 5,400 of 15,113, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Quality-Driven Curation of Remote Sensing Vision-Language Data via Learned Scoring Models2 Mar 2025 0 repositories listed
-
Never too Prim to Swim: An LLM-Enhanced RL-based Adaptive S-Surface Controller for AUVs under Extreme Sea Conditions1 Mar 2025 0 repositories listed
-
Scalable Reinforcement Learning for Virtual Machine Scheduling1 Mar 2025 0 repositories listed
-
Towards Understanding the Benefit of Multitask Representation Learning in Decision Process1 Mar 2025 0 repositories listed
-
Adaptive Reinforcement Learning for State Avoidance in Discrete Event Systems28 Feb 2025 0 repositories listed
-
Hierarchical and Modular Network on Non-prehensile Manipulation in General Environments28 Feb 2025 0 repositories listed
-
Multimodal Dreaming: A Global Workspace Approach to World Model-Based Reinforcement Learning28 Feb 2025 0 repositories listed
-
Subtask-Aware Visual Reward Learning from Segmented Demonstrations28 Feb 2025 0 repositories listed
-
Accelerating Model-Based Reinforcement Learning with State-Space World Models27 Feb 2025 0 repositories listed
-
CarPlanner: Consistent Auto-regressive Trajectory Planning for Large-scale Reinforcement Learning in Autonomous Driving27 Feb 2025 0 repositories listed
-
Improving the Efficiency of a Deep Reinforcement Learning-Based Power Management System for HPC Clusters Using Curriculum Learning27 Feb 2025 0 repositories listed
-
R1-T1: Fully Incentivizing Translation Capability in LLMs via Reasoning Learning27 Feb 2025 0 repositories listed
-
Robust Gymnasium: A Unified Modular Benchmark for Robust Reinforcement Learning27 Feb 2025 0 repositories listed
-
Distill Not Only Data but Also Rewards: Can Smaller Language Models Surpass Larger Ones?26 Feb 2025 0 repositories listed
-
Efficient Reinforcement Learning by Guiding Generalist World Models with Non-Curated Data26 Feb 2025 0 repositories listed
-
Error-related Potential driven Reinforcement Learning for adaptive Brain-Computer Interfaces25 Feb 2025 0 repositories listed
-
FetchBot: Object Fetching in Cluttered Shelves via Zero-Shot Sim2Real25 Feb 2025 0 repositories listed
-
SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution25 Feb 2025 0 repositories listed
-
From Perceptions to Decisions: Wildfire Evacuation Decision Prediction with Behavioral Theory-informed LLMs24 Feb 2025 0 repositories listed
-
Humanoid Whole-Body Locomotion on Narrow Terrain via Dynamic Balance and Reinforcement Learning24 Feb 2025 0 repositories listed
-
Predicting Liquidity-Aware Bond Yields using Causal GANs and Deep Reinforcement Learning with LLM Evaluation24 Feb 2025 0 repositories listed
-
Survey on Strategic Mining in Blockchain: A Reinforcement Learning Approach24 Feb 2025 0 repositories listed
-
Yes, Q-learning Helps Offline In-Context RL24 Feb 2025 0 repositories listed
-
Ensemble RL through Classifier Models: Enhancing Risk-Return Trade-offs in Trading Strategies23 Feb 2025 0 repositories listed
-
Toward Dependency Dynamics in Multi-Agent Reinforcement Learning for Traffic Signal Control23 Feb 2025 0 repositories listed
-
An Autonomous Network Orchestration Framework Integrating Large Language Models with Continual Reinforcement Learning22 Feb 2025 0 repositories listed
-
Together We Rise: Optimizing Real-Time Multi-Robot Task Allocation using Coordinated Heterogeneous Plays22 Feb 2025 0 repositories listed
-
The Evolving Landscape of LLM- and VLM-Integrated Reinforcement Learning21 Feb 2025 0 repositories listed
-
Discovering highly efficient low-weight quantum error-correcting codes with reinforcement learning20 Feb 2025 0 repositories listed
-
Learning from Reward-Free Offline Data: A Case for Planning with Latent Dynamics Models20 Feb 2025 0 repositories listed
-
MLGym: A New Framework and Benchmark for Advancing AI Research Agents20 Feb 2025 0 repositories listed
-
Reinforcement Learning for Ultrasound Image Analysis A Comprehensive Review of Advances and Applications20 Feb 2025 0 repositories listed
-
Reinforcement Learning with Graph Attention for Routing and Wavelength Assignment with Lightpath Reuse20 Feb 2025 0 repositories listed
-
Comprehensive Review on the Control of Heat Pumps for Energy Flexibility in Distribution Networks19 Feb 2025 0 repositories listed
-
Hierarchical RL-MPC for Demand Response Scheduling19 Feb 2025 0 repositories listed
-
Optimizing Gene-Based Testing for Antibiotic Resistance Prediction19 Feb 2025 0 repositories listed
-
SPPD: Self-training with Process Preference Learning Using Dynamic Value Margin19 Feb 2025 0 repositories listed
-
Uncertainty quantification for Markov chains with application to temporal difference learning19 Feb 2025 0 repositories listed
-
A Survey of Sim-to-Real Methods in RL: Progress, Prospects and Challenges with Foundation Models18 Feb 2025 0 repositories listed
-
Demystifying Multilingual Chain-of-Thought in Process Reward Modeling18 Feb 2025 0 repositories listed
-
EPO: Explicit Policy Optimization for Strategic Reasoning in LLMs via Reinforcement Learning18 Feb 2025 0 repositories listed
-
LocalEscaper: A Weakly-supervised Framework with Regional Reconstruction for Scalable Neural TSP Solvers18 Feb 2025 0 repositories listed
-
RAD: Training an End-to-End Driving Policy via Large-Scale 3DGS-based Reinforcement Learning18 Feb 2025 0 repositories listed
-
Addressing Moral Uncertainty using Large Language Models for Ethical Decision-Making17 Feb 2025 0 repositories listed
-
CAMEL: Continuous Action Masking Enabled by Large Language Models for Reinforcement Learning17 Feb 2025 0 repositories listed
-
FitLight: Federated Imitation Learning for Plug-and-Play Autonomous Traffic Signal Control17 Feb 2025 0 repositories listed
-
Hovering Flight of Soft-Actuated Insect-Scale Micro Aerial Vehicles using Deep Reinforcement Learning17 Feb 2025 0 repositories listed
-
Intersectional Fairness in Reinforcement Learning with Large State and Constraint Spaces17 Feb 2025 0 repositories listed
-
Learning Plasma Dynamics and Robust Rampdown Trajectories with Predict-First Experiments at TCV17 Feb 2025 0 repositories listed
-
Robot Deformable Object Manipulation via NMPC-generated Demonstrations in Deep Reinforcement Learning17 Feb 2025 0 repositories listed
-
Scaling Test-Time Compute Without Verification or RL is Suboptimal17 Feb 2025 0 repositories listed
-
FLAG-Trader: Fusion LLM-Agent with Gradient-based Reinforcement Learning for Financial Trading17 Feb 2025 0 repositories listed
-
VLP: Vision-Language Preference Learning for Embodied Manipulation17 Feb 2025 0 repositories listed
-
Scalable Multi-Agent Offline Reinforcement Learning and the Role of Information16 Feb 2025 0 repositories listed
-
Rule-Bottleneck Reinforcement Learning: Joint Explanation and Decision Optimization for Resource Allocation with Language Agents15 Feb 2025 0 repositories listed
-
Tackling the Zero-Shot Reinforcement Learning Loss Directly15 Feb 2025 0 repositories listed
-
BeamDojo: Learning Agile Humanoid Locomotion on Sparse Footholds14 Feb 2025 0 repositories listed
-
Causal Information Prioritization for Efficient Reinforcement Learning14 Feb 2025 0 repositories listed
-
Dynamic Reinforcement Learning for Actors14 Feb 2025 0 repositories listed
-
Provably Efficient RL under Episode-Wise Safety in Constrained MDPs with Linear Function Approximation14 Feb 2025 0 repositories listed
-
Reinforcement Learning in Strategy-Based and Atari Games: A Review of Google DeepMinds Innovations14 Feb 2025 0 repositories listed
-
A Survey of Reinforcement Learning for Optimization in Automation13 Feb 2025 0 repositories listed
-
Diverse Transformer Decoding for Offline Reinforcement Learning Using Financial Algorithmic Approaches13 Feb 2025 0 repositories listed
-
Safe Reinforcement Learning-based Control for Hydrogen Diesel Dual-Fuel Engines13 Feb 2025 0 repositories listed
-
A Real-to-Sim-to-Real Approach to Robotic Manipulation with VLM-Generated Iterative Keypoint Rewards12 Feb 2025 0 repositories listed
-
A Survey on Data-Centric AI: Tabular Learning from Reinforcement Learning and Generative AI Perspective12 Feb 2025 0 repositories listed
-
COMBO-Grasp: Learning Constraint-Based Manipulation for Bimanual Occluded Grasping12 Feb 2025 0 repositories listed
-
12 Feb 2025 0 repositories listed Syntology 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 1 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 1 pointer-only (licence)
-
Necessary and Sufficient Oracles: Toward a Computational Taxonomy For Reinforcement Learning12 Feb 2025 0 repositories listed
-
A Survey of In-Context Reinforcement Learning11 Feb 2025 0 repositories listed
-
Exploratory Diffusion Model for Unsupervised Reinforcement Learning11 Feb 2025 0 repositories listed
-
Model Selection for Off-policy Evaluation: New Algorithms and Experimental Protocol11 Feb 2025 0 repositories listed
-
Near-Optimal Sample Complexity in Reward-Free Kernel-Based Reinforcement Learning11 Feb 2025 0 repositories listed
-
Optimal Actuator Attacks on Autonomous Vehicles Using Reinforcement Learning11 Feb 2025 0 repositories listed
-
Towards a Formal Theory of the Need for Competence via Computational Intrinsic Motivation11 Feb 2025 0 repositories listed
-
Advancing Autonomous VLM Agents via Variational Subgoal-Conditioned Reinforcement Learning11 Feb 2025 0 repositories listed
-
Select before Act: Spatially Decoupled Action Repetition for Continuous Control10 Feb 2025 0 repositories listed
-
Smell of Source: Learning-Based Odor Source Localization with Molecular Communication10 Feb 2025 0 repositories listed
-
Intelligent Offloading in Vehicular Edge Computing: A Comprehensive Review of Deep Reinforcement Learning Approaches and Architectures10 Feb 2025 0 repositories listed
-
Sequential Stochastic Combinatorial Optimization Using Hierarchal Reinforcement Learning8 Feb 2025 0 repositories listed
-
Adversarially-Robust TD Learning with Markovian Data: Finite-Time Rates and Fundamental Limits7 Feb 2025 0 repositories listed
-
Behavior-Regularized Diffusion Policy Optimization for Offline Reinforcement Learning7 Feb 2025 0 repositories listed
-
Convergent NMPC-based Reinforcement Learning Using Deep Expected Sarsa and Nonlinear Temporal Difference Learning7 Feb 2025 0 repositories listed
-
Enhancing Pre-Trained Decision Transformers with Prompt-Tuning Bandits7 Feb 2025 0 repositories listed
-
Learning Strategic Language Agents in the Werewolf Game with Iterative Latent Space Policy Optimization7 Feb 2025 0 repositories listed
-
Towards Smarter Sensing: 2D Clutter Mitigation in RL-Driven Cognitive MIMO Radar7 Feb 2025 0 repositories listed
-
Autotelic Reinforcement Learning: Exploring Intrinsic Motivations for Skill Acquisition in Open-Ended Environments6 Feb 2025 0 repositories listed
-
Behavioral Entropy-Guided Dataset Generation for Offline Reinforcement Learning6 Feb 2025 0 repositories listed
-
Illuminating Spaces: Deep Reinforcement Learning and Laser-Wall Partitioning for Architectural Layout Generation6 Feb 2025 0 repositories listed
-
LLM Alignment as Retriever Optimization: An Information Retrieval Perspective6 Feb 2025 0 repositories listed
-
Mirror Descent Actor Critic via Bounded Advantage Learning6 Feb 2025 0 repositories listed
-
Reinforcement Learning Based Prediction of PID Controller Gains for Quadrotor UAVs6 Feb 2025 0 repositories listed
-
Transforming Multimodal Models into Action Models for Radiotherapy6 Feb 2025 0 repositories listed
-
AI-driven materials design: a mini-review5 Feb 2025 0 repositories listed
-
Calibrated Unsupervised Anomaly Detection in Multivariate Time-series using Reinforcement Learning5 Feb 2025 0 repositories listed
-
OmniRL: In-Context Reinforcement Learning by Large-Scale Meta-Training in Randomized Worlds5 Feb 2025 0 repositories listed
-
Optimizing Electric Vehicles Charging using Large Language Models and Graph Neural Networks5 Feb 2025 0 repositories listed
-
Adviser-Actor-Critic: Eliminating Steady-State Error in Reinforcement Learning Control4 Feb 2025 0 repositories listed
-
Brief analysis of DeepSeek R1 and it's implications for Generative AI4 Feb 2025 0 repositories listed
Syntology lines on 2 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.