Browse State-of-the-Art › Reinforcement Learning (RL) › Papers, page 51
Reinforcement Learning (RL)
Papers archive 2025-07-28
archive papers tagged: 15,113 · with a code link: 4,749 · where Syntology ran a sample: 1,416 (1,186 with a run with no instrument failure, 230 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (1,416 of 15,113 tagged: 1,186 with a run with no instrument failure, 230 where every run was a failure of Syntology's instrument)
Page 51 of 152: papers 5,001 to 5,100 of 15,113, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Of Mice and Machines: A Comparison of Learning Between Real World Mice and RL Agents18 May 2025 0 repositories listed
-
Resolving Latency and Inventory Risk in Market Making with Reinforcement Learning18 May 2025 0 repositories listed
-
UIShift: Enhancing VLM-based GUI Agents through Self-supervised Reinforcement Learning18 May 2025 0 repositories listed
-
AdaCoT: Pareto-Optimal Adaptive Chain-of-Thought Triggering via Reinforcement Learning17 May 2025 0 repositories listed
-
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge17 May 2025 0 repositories listed
-
Online Iterative Self-Alignment for Radiology Report Generation17 May 2025 0 repositories listed
-
Q-Policy: Quantum-Enhanced Policy Evaluation for Scalable Reinforcement Learning17 May 2025 0 repositories listed
-
Solver-Informed RL: Grounding Large Language Models for Authentic Optimization Modeling17 May 2025 0 repositories listed
-
Developing and Integrating Trust Modeling into Multi-Objective Reinforcement Learning for Intelligent Agricultural Management16 May 2025 0 repositories listed
-
Certifying Stability of Reinforcement Learning Policies using Generalized Lyapunov Functions16 May 2025 0 repositories listed
-
ShiQ: Bringing back Bellman to LLMs16 May 2025 0 repositories listed
-
Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes16 May 2025 0 repositories listed
-
Is PRM Necessary? Problem-Solving RL Implicitly Induces PRM Capability in LLMs16 May 2025 0 repositories listed
-
16 May 2025 0 repositories listed Syntology 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Reinforcement Learning for AMR Charging Decisions: The Impact of Reward and Action Space Design16 May 2025 0 repositories listed
-
Spectral Policy Optimization: Coloring your Incorrect Reasoning in GRPO16 May 2025 0 repositories listed
-
Unveiling the Black Box: A Multi-Layer Framework for Explaining Reinforcement Learning-Based Cyber Agents16 May 2025 0 repositories listed
-
Efficient Adaptation of Reinforcement Learning Agents to Sudden Environmental Change15 May 2025 0 repositories listed
-
Fine-tuning Diffusion Policies with Backpropagation Through Diffusion Timesteps15 May 2025 0 repositories listed
-
IN-RIL: Interleaved Reinforcement and Imitation Learning for Policy Fine-Tuning15 May 2025 0 repositories listed
-
Knowledge capture, adaptation and composition (KCAC): A framework for cross-task curriculum learning in robotic manipulation15 May 2025 0 repositories listed
-
Reinforcing the Diffusion Chain of Lateral Thought with Diffusion Language Models15 May 2025 0 repositories listed
-
CEC-Zero: Chinese Error Correction Solution Based on LLM14 May 2025 0 repositories listed
-
Reinforcement Learning for Individual Optimal Policy from Heterogeneous Data14 May 2025 0 repositories listed
-
Risk-Aware Safe Reinforcement Learning for Control of Stochastic Linear Systems14 May 2025 0 repositories listed
-
14 May 2025 0 repositories listed Syntology 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 9 harvested samples) · 1 pointer-only (licence)
-
Adaptive Diffusion Policy Optimization for Robotic Manipulation13 May 2025 0 repositories listed
-
Adaptive Security Policy Management in Cloud Environments Using Reinforcement Learning13 May 2025 0 repositories listed
-
Automatic Curriculum Learning for Driving Scenarios: Towards Robust and Efficient Reinforcement Learning13 May 2025 0 repositories listed
-
DSADF: Thinking Fast and Slow for Decision Making13 May 2025 0 repositories listed
-
Generalization in Monitored Markov Decision Processes (Mon-MDPs)13 May 2025 0 repositories listed
-
Modeling Unseen Environments with Language-guided Composable Causal Components in Reinforcement Learning13 May 2025 0 repositories listed
-
Preference Optimization for Combinatorial Optimization Problems13 May 2025 0 repositories listed
-
Reinforcement Learning-based Fault-Tolerant Control for Quadrotor with Online Transformer Adaptation13 May 2025 0 repositories listed
-
Scaling Multi Agent Reinforcement Learning for Underwater Acoustic Tracking via Autonomous Vehicles13 May 2025 0 repositories listed
-
Cache-Efficient Posterior Sampling for Reinforcement Learning with LLM-Derived Priors Across Discrete and Continuous Domains12 May 2025 0 repositories listed
-
Combining Bayesian Inference and Reinforcement Learning for Agent Decision Making: A Review12 May 2025 0 repositories listed
-
DARLR: Dual-Agent Offline Reinforcement Learning for Recommender Systems with Dynamic Reward12 May 2025 0 repositories listed
-
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning12 May 2025 0 repositories listed
-
INTELLECT-2: A Reasoning Model Trained Through Globally Decentralized Reinforcement Learning12 May 2025 0 repositories listed
-
The Exploratory Multi-Asset Mean-Variance Portfolio Selection using Reinforcement Learning12 May 2025 0 repositories listed
-
Design and Experimental Test of Datatic Approximate Optimal Filter in Nonlinear Dynamic Systems11 May 2025 0 repositories listed
-
FACET: Force-Adaptive Control via Impedance Reference Tracking for Legged Robots11 May 2025 0 repositories listed
-
Learning Value of Information towards Joint Communication and Control in 6G V2X11 May 2025 0 repositories listed
-
Reinforcement Learning (RL) Meets Urban Climate Modeling: Investigating the Efficacy and Impacts of RL-Based HVAC Control11 May 2025 0 repositories listed
-
X-Sim: Cross-Embodiment Learning via Real-to-Sim-to-Real11 May 2025 0 repositories listed
-
Balancing Progress and Safety: A Novel Risk-Aware Objective for RL in Autonomous Driving10 May 2025 0 repositories listed
-
REFINE-AF: A Task-Agnostic Framework to Align Language Models via Self-Generated Instructions using Reinforcement Learning from Automated Feedback10 May 2025 0 repositories listed
-
Video-Enhanced Offline Reinforcement Learning: A Model-Based Approach10 May 2025 0 repositories listed
-
Active Perception for Tactile Sensing: A Task-Agnostic Attention-Based Approach9 May 2025 0 repositories listed
-
Interaction-Aware Parameter Privacy-Preserving Data Sharing in Coupled Systems via Particle Filter Reinforcement Learning9 May 2025 0 repositories listed
-
Pretraining a Shared Q-Network for Data-Efficient Offline Reinforcement Learning9 May 2025 0 repositories listed
-
Remote Rowhammer Attack using Adversarial Observations on Federated Learning Clients9 May 2025 0 repositories listed
-
Enhancing Reinforcement Learning for the Floorplanning of Analog ICs with Beam Search8 May 2025 0 repositories listed
-
Multi-agent Embodied AI: Advances and Future Directions8 May 2025 0 repositories listed
-
On Corruption-Robustness in Performative Reinforcement Learning8 May 2025 0 repositories listed
-
Reinforcement Learning for Game-Theoretic Resource Allocation on Graphs8 May 2025 0 repositories listed
-
RL-DAUNCE: Reinforcement Learning-Driven Data Assimilation with Uncertainty-Aware Constrained Ensembles8 May 2025 0 repositories listed
-
Taming OOD Actions for Offline Reinforcement Learning: An Advantage-Based Approach8 May 2025 0 repositories listed
-
USPR: Learning a Unified Solver for Profiled Routing8 May 2025 0 repositories listed
-
Extending a Quantum Reinforcement Learning Exploration Policy with Flags to Connect Four7 May 2025 0 repositories listed
-
Fight Fire with Fire: Defending Against Malicious RL Fine-Tuning via Reward Neutralization7 May 2025 0 repositories listed
-
Putting the Value Back in RL: Better Test-Time Scaling by Unifying LLM Reasoners With Verifiers7 May 2025 0 repositories listed
-
Risk-sensitive Reinforcement Learning Based on Convex Scoring Functions7 May 2025 0 repositories listed
-
Actor-Critics Can Achieve Optimal Sample Efficiency6 May 2025 0 repositories listed
-
AMO: Adaptive Motion Optimization for Hyper-Dexterous Humanoid Whole-Body Control6 May 2025 0 repositories listed
-
Decentralized Distributed Proximal Policy Optimization (DD-PPO) for High Performance Computing Scheduling on Multi-User Systems6 May 2025 0 repositories listed
-
Deep Q-Network (DQN) multi-agent reinforcement learning (MARL) for Stock Trading6 May 2025 0 repositories listed
-
The Steganographic Potentials of Language Models6 May 2025 0 repositories listed
-
VLM Q-Learning: Aligning Vision-Language Models for Interactive Decision-Making6 May 2025 0 repositories listed
-
Automated Hybrid Reward Scheduling via Large Language Models for Robotic Skill Learning5 May 2025 0 repositories listed
-
Online Phase Estimation of Human Oscillatory Motions using Deep Learning5 May 2025 0 repositories listed
-
Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study4 May 2025 0 repositories listed
-
Prompt-responsive Object Retrieval with Memory-augmented Student-Teacher Learning4 May 2025 0 repositories listed
-
Analytic Energy-Guided Policy Optimization for Offline Reinforcement Learning3 May 2025 0 repositories listed
-
World Model-Based Learning for Long-Term Age of Information Minimization in Vehicular Networks3 May 2025 0 repositories listed
-
Stabilizing Temporal Difference Learning via Implicit Stochastic Recursion2 May 2025 0 repositories listed
-
A General Approach of Automated Environment Design for Learning the Optimal Power Flow1 May 2025 0 repositories listed
-
Implicit Neural-Representation Learning for Elastic Deformable-Object Manipulations1 May 2025 0 repositories listed
-
Leveraging Partial SMILES Validation Scheme for Enhanced Drug Design in Reinforcement Learning Frameworks1 May 2025 0 repositories listed
-
MULE: Multi-terrain and Unknown Load Adaptation for Effective Quadrupedal Locomotion1 May 2025 0 repositories listed
-
Adaptive 3D UI Placement in Mixed Reality Using Deep Reinforcement Learning30 Apr 2025 0 repositories listed
-
Phi-4-Mini-Reasoning: Exploring the Limits of Small Reasoning Language Models in Math30 Apr 2025 0 repositories listed
-
Phi-4-reasoning Technical Report30 Apr 2025 0 repositories listed
-
Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models30 Apr 2025 0 repositories listed
-
A Survey on GUI Agents with Foundation Models Enhanced by Reinforcement Learning29 Apr 2025 0 repositories listed
-
PRISM: Projection-based Reward Integration for Scene-Aware Real-to-Sim-to-Real Transfer with Few Demonstrations29 Apr 2025 0 repositories listed
-
Token-Efficient RL for LLM Reasoning29 Apr 2025 0 repositories listed
-
AI Recommendation Systems for Lane-Changing Using Adherence-Aware Reinforcement Learning28 Apr 2025 0 repositories listed
-
An Automated Reinforcement Learning Reward Design Framework with Large Language Model for Cooperative Platoon Coordination28 Apr 2025 0 repositories listed
-
Interactive Double Deep Q-network: Integrating Human Interventions and Evaluative Predictions in Reinforcement Learning of Autonomous Driving28 Apr 2025 0 repositories listed
-
Reinforcement Learning-Based Heterogeneous Multi-Task Optimization in Semantic Broadcast Communications28 Apr 2025 0 repositories listed
-
LLMs for Engineering: Teaching Models to Design High Powered Rockets27 Apr 2025 0 repositories listed
-
Depth-Constrained ASV Navigation with Deep RL and Limited Sensing25 Apr 2025 0 repositories listed
-
Explainable AI for UAV Mobility Management: A Deep Q-Network Approach for Handover Minimization25 Apr 2025 0 repositories listed
-
LLM-hRIC: LLM-empowered Hierarchical RAN Intelligent Control for O-RAN25 Apr 2025 0 repositories listed
-
Integrating Learning-Based Manipulation and Physics-Based Locomotion for Whole-Body Badminton Robot Control24 Apr 2025 0 repositories listed
-
SAPO-RL: Sequential Actuator Placement Optimization for Fuselage Assembly via Reinforcement Learning24 Apr 2025 0 repositories listed
-
Training Large Language Models to Reason via EM Policy Gradient24 Apr 2025 0 repositories listed
-
Data-Assimilated Model-Based Reinforcement Learning for Partially Observed Chaotic Flows23 Apr 2025 0 repositories listed
Syntology lines on 2 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.