Browse State-of-the-Art › Reinforcement Learning › Papers, page 44
Reinforcement Learning
Papers archive 2025-07-28
archive papers tagged: 13,178 · with a code link: 4,183 · where Syntology ran a sample: 1,175 (988 with a run with no instrument failure, 187 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (1,175 of 13,178 tagged: 988 with a run with no instrument failure, 187 where every run was a failure of Syntology's instrument)
Page 44 of 132: papers 4,301 to 4,400 of 13,178, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Interleaved Reasoning for Large Language Models via Reinforcement Learning26 May 2025 0 repositories listed
-
Outcome-Based Online Reinforcement Learning: Algorithms and Fundamental Limits26 May 2025 0 repositories listed
-
TeViR: Text-to-Video Reward with Diffusion Models for Efficient Reinforcement Learning26 May 2025 0 repositories listed
-
The challenge of hidden gifts in multi-agent reinforcement learning26 May 2025 0 repositories listed
-
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection26 May 2025 0 repositories listed
-
Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning25 May 2025 0 repositories listed
-
Semi-pessimistic Reinforcement Learning25 May 2025 0 repositories listed
-
EdgeAgentX: A Novel Framework for Agentic AI at the Edge in Military Communication Networks24 May 2025 0 repositories listed
-
G1: Teaching LLMs to Reason on Graphs with Reinforcement Learning24 May 2025 0 repositories listed
-
GenPO: Generative Diffusion Models Meet On-Policy Reinforcement Learning24 May 2025 0 repositories listed
-
Pessimism Principle Can Be Effective: Towards a Framework for Zero-Shot Transfer Reinforcement Learning24 May 2025 0 repositories listed
-
Diffusion Self-Weighted Guidance for Offline Reinforcement Learning23 May 2025 0 repositories listed
-
Outcome-based Reinforcement Learning to Predict the Future23 May 2025 0 repositories listed
-
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning22 May 2025 0 repositories listed
-
Backdoors in DRL: Four Environments Focusing on In-distribution Triggers22 May 2025 0 repositories listed
-
Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only22 May 2025 0 repositories listed
-
Incentivizing Dual Process Thinking for Efficient Large Language Model Reasoning22 May 2025 0 repositories listed
-
ManipLVM-R1: Reinforcement Learning for Reasoning in Embodied Manipulation with Large Vision-Language Models22 May 2025 0 repositories listed
-
Meta-reinforcement learning with minimum attention22 May 2025 0 repositories listed
-
Offline Guarded Safe Reinforcement Learning for Medical Treatment Optimization Strategies22 May 2025 0 repositories listed
-
22 May 2025 0 repositories listed
-
Reinforcement Learning for Stock Transactions22 May 2025 0 repositories listed
-
Reward-Aware Proto-Representations in Reinforcement Learning22 May 2025 0 repositories listed
-
Risk-Averse Reinforcement Learning with Itakura-Saito Loss22 May 2025 0 repositories listed
-
AM-PPO: (Advantage) Alpha-Modulation with Proximal Policy Optimization21 May 2025 0 repositories listed
-
Filtering Learning Histories Enhances In-Context Reinforcement Learning21 May 2025 0 repositories listed
-
Learning from Algorithm Feedback: One-Shot SAT Solver Guidance with GNNs21 May 2025 0 repositories listed
-
Pass@K Policy Optimization: Solving Harder Reinforcement Learning Problems21 May 2025 0 repositories listed
-
Reverse Engineering Human Preferences with Reinforcement Learning21 May 2025 0 repositories listed
-
World Models as Reference Trajectories for Rapid Motor Adaptation21 May 2025 0 repositories listed
-
Bellman operator convergence enhancements in reinforcement learning algorithms20 May 2025 0 repositories listed
-
Embedded Mean Field Reinforcement Learning for Perimeter-defense Game20 May 2025 0 repositories listed
-
Energy-Efficient Deep Reinforcement Learning with Spiking Transformers20 May 2025 0 repositories listed
-
FlowQ: Energy-Guided Flow Policies for Offline Reinforcement Learning20 May 2025 0 repositories listed
-
Interpretable Reinforcement Learning for Load Balancing using Kolmogorov-Arnold Networks20 May 2025 0 repositories listed
-
Reinforcement Learning from User Feedback20 May 2025 0 repositories listed
-
SHARP: Synthesizing High-quality Aligned Reasoning Problems for Large Reasoning Models Reinforcement Learning20 May 2025 0 repositories listed
-
Normalized Cut with Reinforcement Learning in Constrained Action Space20 May 2025 0 repositories listed
-
Time Reversal Symmetry for Efficient Robotic Manipulations in Deep Reinforcement Learning20 May 2025 0 repositories listed
-
Visionary-R1: Mitigating Shortcuts in Visual Reasoning with Reinforcement Learning20 May 2025 0 repositories listed
-
Composing Dextrous Grasping and In-hand Manipulation via Scoring with a Reinforcement Learning Critic19 May 2025 0 repositories listed
-
Dynamic Sight Range Selection in Multi-Agent Reinforcement Learning19 May 2025 0 repositories listed
-
Multi-parameter Control for the (1+(λ,λ))-GA on OneMax via Deep Reinforcement Learning19 May 2025 0 repositories listed
-
When a Reinforcement Learning Agent Encounters Unknown Unknowns19 May 2025 0 repositories listed
-
A universal policy wrapper with guarantees18 May 2025 0 repositories listed
-
Imagination-Limited Q-Learning for Offline Reinforcement Learning18 May 2025 0 repositories listed
-
Table-R1: Region-based Reinforcement Learning for Table Understanding18 May 2025 0 repositories listed
-
Solver-Informed RL: Grounding Large Language Models for Authentic Optimization Modeling17 May 2025 0 repositories listed
-
Deep Symbolic Optimization: Reinforcement Learning for Symbolic Mathematics16 May 2025 0 repositories listed
-
Prior-Guided Diffusion Planning for Offline Reinforcement Learning16 May 2025 0 repositories listed
-
Improving Assembly Code Performance with Large Language Models via Reinforcement Learning16 May 2025 0 repositories listed
-
16 May 2025 0 repositories listed Syntology 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Approximated Behavioral Metric-based State Projection for Federated Reinforcement Learning15 May 2025 0 repositories listed
-
Efficient Adaptation of Reinforcement Learning Agents to Sudden Environmental Change15 May 2025 0 repositories listed
-
Fixing Incomplete Value Function Decomposition for Multi-Agent Reinforcement Learning15 May 2025 0 repositories listed
-
Offline Reinforcement Learning for Microgrid Voltage Regulation15 May 2025 0 repositories listed
-
ORL-LDM: Offline Reinforcement Learning Guided Latent Diffusion Model Super-Resolution Reconstruction15 May 2025 0 repositories listed
-
Sample Complexity of Distributionally Robust Average-Reward Reinforcement Learning15 May 2025 0 repositories listed
-
Reinforcement Learning for Individual Optimal Policy from Heterogeneous Data14 May 2025 0 repositories listed
-
A Practical Introduction to Deep Reinforcement Learning13 May 2025 0 repositories listed
-
Adaptive Security Policy Management in Cloud Environments Using Reinforcement Learning13 May 2025 0 repositories listed
-
Deep Reinforcement Learning for Power Grid Multi-Stage Cascading Failure Mitigation13 May 2025 0 repositories listed
-
Enhancing Aerial Combat Tactics through Hierarchical Multi-Agent Reinforcement Learning13 May 2025 0 repositories listed
-
A Theoretical Framework for Explaining Reinforcement Learning with Shapley Values12 May 2025 0 repositories listed
-
Explainable Reinforcement Learning Agents Using World Models12 May 2025 0 repositories listed
-
INTELLECT-2: A Reasoning Model Trained Through Globally Decentralized Reinforcement Learning12 May 2025 0 repositories listed
-
Multi-Objective Reinforcement Learning for Energy-Efficient Industrial Control12 May 2025 0 repositories listed
-
Multi-source Plume Tracing via Multi-Agent Reinforcement Learning12 May 2025 0 repositories listed
-
Online Episodic Convex Reinforcement Learning12 May 2025 0 repositories listed
-
S-GRPO: Early Exit via Reinforcement Learning in Reasoning Models12 May 2025 0 repositories listed
-
Self Rewarding Self Improving12 May 2025 0 repositories listed
-
SEM: Reinforcement Learning for Search-Efficient Large Language Models12 May 2025 0 repositories listed
-
What Matters for Batch Online Reinforcement Learning in Robotics?12 May 2025 0 repositories listed
-
Video-Enhanced Offline Reinforcement Learning: A Model-Based Approach10 May 2025 0 repositories listed
-
Offline Multi-agent Reinforcement Learning via Score Decomposition9 May 2025 0 repositories listed
-
Morphologically Symmetric Reinforcement Learning for Ambidextrous Bimanual Manipulation8 May 2025 0 repositories listed
-
Multi-Objective Reinforcement Learning for Adaptive Personalized Autonomous Driving8 May 2025 0 repositories listed
-
On Corruption-Robustness in Performative Reinforcement Learning8 May 2025 0 repositories listed
-
Reasoning Models Don't Always Say What They Think8 May 2025 0 repositories listed
-
Reinforcement Learning for Game-Theoretic Resource Allocation on Graphs8 May 2025 0 repositories listed
-
Is there Value in Reinforcement Learning?7 May 2025 0 repositories listed
-
Optimization of Infectious Disease Intervention Measures Based on Reinforcement Learning - Empirical analysis based on UK COVID-19 epidemic data7 May 2025 0 repositories listed
-
Risk-sensitive Reinforcement Learning Based on Convex Scoring Functions7 May 2025 0 repositories listed
-
Trajectory Entropy Reinforcement Learning for Predictable and Robust Control7 May 2025 0 repositories listed
-
DYSTIL: Dynamic Strategy Induction with Large Language Models for Reinforcement Learning6 May 2025 0 repositories listed
-
Interpretable Learning Dynamics in Unsupervised Reinforcement Learning6 May 2025 0 repositories listed
-
Null Counterfactual Factor Interactions for Goal-Conditioned Reinforcement Learning6 May 2025 0 repositories listed
-
Policy-labeled Preference Learning: Is Preference Enough for RLHF?6 May 2025 0 repositories listed
-
A Goal-Oriented Reinforcement Learning-Based Path Planning Algorithm for Modular Self-Reconfigurable Satellites4 May 2025 0 repositories listed
-
A Synergistic Framework of Nonlinear Acoustic Computing and Reinforcement Learning for Real-World Human-Robot Interaction4 May 2025 0 repositories listed
-
Training Environment for High Performance Reinforcement Learning4 May 2025 0 repositories listed
-
Analytic Energy-Guided Policy Optimization for Offline Reinforcement Learning3 May 2025 0 repositories listed
-
Modeling Behavioral Preferences of Cyber Adversaries Using Inverse Reinforcement Learning2 May 2025 0 repositories listed
-
Skill-based Safe Reinforcement Learning with Risk Planning2 May 2025 0 repositories listed
-
Emergence of Roles in Robotic Teams with Model Sharing and Limited Communication1 May 2025 0 repositories listed
-
Intelligent Task Scheduling for Microservices via A3C-Based Reinforcement Learning1 May 2025 0 repositories listed
-
Reinforcement Learning with Continuous Actions Under Unmeasured Confounding1 May 2025 0 repositories listed
-
Variational OOD State Correction for Offline Reinforcement Learning1 May 2025 0 repositories listed
-
Intelligent Task Offloading in VANETs: A Hybrid AI-Driven Approach for Low-Latency and Energy Efficiency29 Apr 2025 0 repositories listed
-
Quantum-Enhanced Hybrid Reinforcement Learning Framework for Dynamic Path Planning in Autonomous Systems29 Apr 2025 0 repositories listed
Syntology lines on 1 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.