Browse State-of-the-Art › Reinforcement Learning (RL) › Papers, page 49
Reinforcement Learning (RL)
Papers archive 2025-07-28
archive papers tagged: 15,113 · with a code link: 4,749 · where Syntology ran a sample: 1,416 (1,186 with a run with no instrument failure, 230 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (1,416 of 15,113 tagged: 1,186 with a run with no instrument failure, 230 where every run was a failure of Syntology's instrument)
Page 49 of 152: papers 4,801 to 4,900 of 15,113, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
IntelliLung: Advancing Safe Mechanical Ventilation using Offline RL with Hybrid Actions and Clinically Aligned Rewards17 Jun 2025 0 repositories listed
-
PeRL: Permutation-Enhanced Reinforcement Learning for Interleaved Vision-Language Reasoning17 Jun 2025 0 repositories listed
-
Reasoning with Exploration: An Entropy Perspective17 Jun 2025 0 repositories listed
-
Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs17 Jun 2025 0 repositories listed
-
Unsupervised Skill Discovery through Skill Regions Differentiation17 Jun 2025 0 repositories listed
-
Zeroth-Order Optimization is Secretly Single-Step Policy Optimization17 Jun 2025 0 repositories listed
-
A Technical Study into Small Reasoning Language Models16 Jun 2025 0 repositories listed
-
AceReason-Nemotron 1.1: Advancing Math and Code Reasoning through SFT and RL Synergy16 Jun 2025 0 repositories listed
-
Can you see how I learn? Human observers' inferences about Reinforcement Learning agents' learning processes16 Jun 2025 0 repositories listed
-
Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning16 Jun 2025 0 repositories listed
-
ReinDSplit: Reinforced Dynamic Split Learning for Pest Recognition in Precision Agriculture16 Jun 2025 0 repositories listed
-
RL-Guided MPC for Autonomous Greenhouse Control16 Jun 2025 0 repositories listed
-
Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation16 Jun 2025 0 repositories listed
-
StaQ it! Growing neural networks for Policy Mirror Descent16 Jun 2025 0 repositories listed
-
16 Jun 2025 0 repositories listed Syntology 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
CAPO: Reinforcing Consistent Reasoning in Medical Decision-Making15 Jun 2025 0 repositories listed
-
Federated Neuroevolution O-RAN: Enhancing the Robustness of Deep Reinforcement Learning xApps15 Jun 2025 0 repositories listed
-
MM-R5: MultiModal Reasoning-Enhanced ReRanker via Reinforcement Learning for Document Retrieval14 Jun 2025 0 repositories listed
-
Automated Treatment Planning for Interstitial HDR Brachytherapy for Locally Advanced Cervical Cancer using Deep Reinforcement Learning13 Jun 2025 0 repositories listed
-
Eliciting Reasoning in Language Models with Cognitive Tools13 Jun 2025 0 repositories listed
-
LearnAlign: Reasoning Data Selection for Reinforcement Learning in Large Language Models Based on Improved Gradient Alignment13 Jun 2025 0 repositories listed
-
ReVeal: Self-Evolving Code Agents via Iterative Generation-Verification13 Jun 2025 0 repositories listed
-
Magistral12 Jun 2025 0 repositories listed
-
PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier12 Jun 2025 0 repositories listed
-
Attention on flow control: transformer-based reinforcement learning for lift regulation in highly disturbed flows11 Jun 2025 0 repositories listed
-
A Survey on the Role of Artificial Intelligence and Machine Learning in 6G-V2X Applications11 Jun 2025 0 repositories listed
-
Automatic Treatment Planning using Reinforcement Learning for High-dose-rate Prostate Brachytherapy11 Jun 2025 0 repositories listed
-
Bridging Continuous-time LQR and Reinforcement Learning via Gradient Flow of the Bellman Error11 Jun 2025 0 repositories listed
-
How to Provably Improve Return Conditioned Supervised Learning?10 Jun 2025 0 repositories listed
-
MasHost Builds It All: Autonomous Multi-Agent System Directed by Reinforcement Learning10 Jun 2025 0 repositories listed
-
Robust Evolutionary Multi-Objective Network Architecture Search for Reinforcement Learning (EMNAS-RL)10 Jun 2025 0 repositories listed
-
DeepForm: Reasoning Large Language Model for Communication System Formulation10 Jun 2025 0 repositories listed
-
Exploration by Random Reward Perturbation10 Jun 2025 0 repositories listed
-
Optimal Operating Strategy for PV-BESS Households: Balancing Self-Consumption and Self-Sufficiency10 Jun 2025 0 repositories listed
-
Policy-Based Trajectory Clustering in Offline Reinforcement Learning10 Jun 2025 0 repositories listed
-
TGRPO :Fine-tuning Vision-Language-Action Model via Trajectory-wise Group Relative Policy Optimization10 Jun 2025 0 repositories listed
-
AbstRaL: Augmenting LLMs' Reasoning by Reinforcing Abstract Thinking9 Jun 2025 0 repositories listed
-
Bingo: Boosting Efficient Reasoning of LLMs via Dynamic and Significance-based Reinforcement Learning9 Jun 2025 0 repositories listed
-
Decentralizing Multi-Agent Reinforcement Learning with Temporal Causal Information9 Jun 2025 0 repositories listed
-
DeepVideo-R1: Video Reinforcement Fine-Tuning via Difficulty-aware Regressive GRPO9 Jun 2025 0 repositories listed
-
LUCIFER: Language Understanding and Context-Infused Framework for Exploration and Behavior Refinement9 Jun 2025 0 repositories listed
-
Reinforcement Pre-Training9 Jun 2025 0 repositories listed
-
Through the Valley: Path to Effective Long CoT Training for Small Language Models9 Jun 2025 0 repositories listed
-
CARoL: Context-aware Adaptation for Robot Learning8 Jun 2025 0 repositories listed
-
Learning to Clarify by Reinforcement Learning Through Reward-Weighted Fine-Tuning8 Jun 2025 0 repositories listed
-
On the Generalization of Data-Assisted Control in port-Hamiltonian Systems (DAC-pH)8 Jun 2025 0 repositories listed
-
QForce-RL: Quantized FPGA-Optimized Reinforcement Learning Compute Engine8 Jun 2025 0 repositories listed
-
Reliable Critics: Monotonic Improvement and Convergence Guarantees for Reinforcement Learning8 Jun 2025 0 repositories listed
-
Safety-Aware Reinforcement Learning for Control via Risk-Sensitive Action-Value Iteration and Quantile Regression8 Jun 2025 0 repositories listed
-
CodeContests+: High-Quality Test Case Generation for Competitive Programming6 Jun 2025 0 repositories listed
-
Prompting Wireless Networks: Reinforced In-Context Learning for Power Control6 Jun 2025 0 repositories listed
-
Towards Infant Sleep-Optimized Driving: Synergizing Wearable and Vehicle Sensing in Intelligent Cruise Control6 Jun 2025 0 repositories listed
-
Beyond Accuracy: Dissecting Mathematical Reasoning for LLMs Under Reinforcement Learning5 Jun 2025 0 repositories listed
-
Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models5 Jun 2025 0 repositories listed
-
On the Mechanism of Reasoning Pattern Selection in Reinforcement Learning for Language Models5 Jun 2025 0 repositories listed
-
Regret-Optimal Q-Learning with Low Cost for Single-Agent and Federated Reinforcement Learning5 Jun 2025 0 repositories listed
-
Safe Planning and Policy Optimization via World Model Learning5 Jun 2025 0 repositories listed
-
A Lyapunov Drift-Plus-Penalty Method Tailored for Reinforcement Learning with Queue Stability4 Jun 2025 0 repositories listed
-
Advancing Multimodal Reasoning: From Optimized Cold Start to Staged Reinforcement Learning4 Jun 2025 0 repositories listed
-
CORE: Constraint-Aware One-Step Reinforcement Learning for Simulation-Guided Neural Network Accelerator Design4 Jun 2025 0 repositories listed
-
Learning-at-Criticality in Large Language Models for Quantum Field Theory and Beyond4 Jun 2025 0 repositories listed
-
SLAC: Simulation-Pretrained Latent Action Space for Whole-Body Real-World RL4 Jun 2025 0 repositories listed
-
Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback3 Jun 2025 0 repositories listed
-
Joint Modeling for Learning Decision-Making Dynamics in Behavioral Experiments3 Jun 2025 0 repositories listed
-
Learned Controllers for Agile Quadrotors in Pursuit-Evasion Games3 Jun 2025 0 repositories listed
-
Unleashing the Reasoning Potential of Pre-trained LLMs by Critique Fine-Tuning on One Problem3 Jun 2025 0 repositories listed
-
Data-assimilated model-informed reinforcement learning2 Jun 2025 0 repositories listed
-
KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning2 Jun 2025 0 repositories listed
-
Knowledge or Reasoning? A Close Look at How LLMs Think Across Domains2 Jun 2025 0 repositories listed
-
SRPO: Enhancing Multimodal LLM Reasoning via Reflection-Aware Reinforcement Learning2 Jun 2025 0 repositories listed
-
Trajectory First: A Curriculum for Discovering Diverse Policies2 Jun 2025 0 repositories listed
-
A Reinforcement Learning Approach for RIS-aided Fair Communications1 Jun 2025 0 repositories listed
-
DriveMind: A Dual-VLM based Reinforcement Learning Framework for Autonomous Driving1 Jun 2025 0 repositories listed
-
ARIA: Training Language Agents with Intention-Driven Reward Aggregation31 May 2025 0 repositories listed
-
MMedAgent-RL: Optimizing Multi-Agent Collaboration for Multimodal Medical Reasoning31 May 2025 0 repositories listed
-
Reinforcement Learning for Hanabi31 May 2025 0 repositories listed
-
Balancing Profit and Fairness in Risk-Based Pricing Markets30 May 2025 0 repositories listed
-
How Much Backtracking is Enough? Exploring the Interplay of SFT and RL in Enhancing LLM Reasoning30 May 2025 0 repositories listed
-
Pangu DeepDiver: Adaptive Search Intensity Scaling via Open-Web Reinforcement Learning30 May 2025 0 repositories listed
-
Proxy Target: Bridging the Gap Between Discrete Spiking Neural Networks and Continuous Control30 May 2025 0 repositories listed
-
Reason-SVG: Hybrid Reward RL for Aha-Moments in Vector Graphics Generation30 May 2025 0 repositories listed
-
ROAD: Responsibility-Oriented Reward Design for Reinforcement Learning in Autonomous Driving30 May 2025 0 repositories listed
-
ADG: Ambient Diffusion-Guided Dataset Recovery for Corruption-Robust Offline Reinforcement Learning29 May 2025 0 repositories listed
-
Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization29 May 2025 0 repositories listed
-
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners29 May 2025 0 repositories listed
-
Composite Flow Matching for Reinforcement Learning with Shifted-Dynamics Data29 May 2025 0 repositories listed
-
Contextual Integrity in LLMs via Reasoning and Reinforcement Learning29 May 2025 0 repositories listed
-
DIP-R1: Deep Inspection and Perception with RL Looking Through and Understanding Complex Scenes29 May 2025 0 repositories listed
-
Diversity-Aware Policy Optimization for Large Language Model Reasoning29 May 2025 0 repositories listed
-
Fine-Tuning Next-Scale Visual Autoregressive Models with Group Relative Policy Optimization29 May 2025 0 repositories listed
-
Fortune: Formula-Driven Reinforcement Learning for Symbolic Table Reasoning in Language Models29 May 2025 0 repositories listed
-
Grower-in-the-Loop Interactive Reinforcement Learning for Greenhouse Climate Control29 May 2025 0 repositories listed
-
Hybrid Cross-domain Robust Reinforcement Learning29 May 2025 0 repositories listed
-
Let's Reason Formally: Natural-Formal Hybrid Reasoning Enhances LLM's Math Capability29 May 2025 0 repositories listed
-
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Trainin29 May 2025 0 repositories listed
-
Measure gradients, not activations! Enhancing neuronal activity in deep reinforcement learning29 May 2025 0 repositories listed
-
Reinforcement Learning for Better Verbalized Confidence in Long-Form Generation29 May 2025 0 repositories listed
-
Unsupervised Transcript-assisted Video Summarization and Highlight Detection29 May 2025 0 repositories listed
-
A Provable Approach for End-to-End Safe Reinforcement Learning28 May 2025 0 repositories listed
Syntology lines on 2 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.