Browse State-of-the-Art › reinforcement-learning › Papers, page 42
reinforcement-learning
Papers archive 2025-07-28
archive papers tagged: 13,427 · with a code link: 4,119 · where Syntology ran a sample: 1,165 (973 with a run with no instrument failure, 192 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (1,165 of 13,427 tagged: 973 with a run with no instrument failure, 192 where every run was a failure of Syntology's instrument)
Page 42 of 135: papers 4,101 to 4,200 of 13,427, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
25 Nov 2015 1 repository listed
-
19 Nov 2015 1 repository listed
-
19 Nov 2015 1 repository listed
-
21 Sep 2015 1 repository listed
-
3 Jul 2015 1 repository listed
-
4 May 2015 1 repository listed
-
22 Jun 2014 1 repository listed
-
30 Apr 2014 1 repository listed
-
15 Apr 2014 1 repository listed
-
4 Apr 2014 1 repository listed
-
4 Feb 2014 1 repository listed
-
1 Dec 2013 1 repository listed
-
18 Jun 2012 1 repository listed
-
1 Dec 2011 1 repository listed
-
1 Jul 2008 1 repository listed
-
4 Dec 2003 1 repository listed
-
6 Aug 1999 1 repository listed
-
1 May 1992 1 repository listed
-
CUDA-L1: Improving CUDA Optimization via Contrastive Reinforcement Learning18 Jul 2025 0 repositories listed
-
Aligning Humans and Robots via Reinforcement Learning from Implicit Human Feedback17 Jul 2025 0 repositories listed
-
Autonomous Resource Management in Microservice Systems via Reinforcement Learning17 Jul 2025 0 repositories listed
-
From Novelty to Imitation: Self-Distilled Rewards for Offline Reinforcement Learning17 Jul 2025 0 repositories listed
-
Spectral Bellman Method: Unifying Representation and Exploration in RL17 Jul 2025 0 repositories listed
-
VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning17 Jul 2025 0 repositories listed
-
A Survey of Explainable Reinforcement Learning: Targets, Methods and Needs16 Jul 2025 0 repositories listed
-
Distributional Reinforcement Learning on Path-dependent Options16 Jul 2025 0 repositories listed
-
Improving Reinforcement Learning Sample-Efficiency using Local Approximation16 Jul 2025 0 repositories listed
-
Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training16 Jul 2025 0 repositories listed
-
Thought Purity: Defense Paradigm For Chain-of-Thought Attack16 Jul 2025 0 repositories listed
-
Functional Emotion Modeling in Biomimetic Reinforcement Learning15 Jul 2025 0 repositories listed
-
Local Pairwise Distance Matching for Backpropagation-Free Reinforcement Learning15 Jul 2025 0 repositories listed
-
Tactical Decision for Multi-UGV Confrontation with a Vision-Language Model-Based Commander15 Jul 2025 0 repositories listed
-
Continual Reinforcement Learning by Planning with Online World Models12 Jul 2025 0 repositories listed
-
CTRLS: Chain-of-Thought Reasoning via Latent State-Transition10 Jul 2025 0 repositories listed
-
From Curiosity to Competence: How World Models Interact with the Dynamics of Exploration10 Jul 2025 0 repositories listed
-
Artificial Generals Intelligence: Mastering Generals.io with Reinforcement Learning9 Jul 2025 0 repositories listed
-
2048: Reinforcement Learning in a Delayed Reward Environment7 Jul 2025 0 repositories listed
-
Epistemically-guided forward-backward exploration7 Jul 2025 0 repositories listed
-
Learn Globally, Speak Locally: Bridging the Gaps in Multilingual Reasoning7 Jul 2025 0 repositories listed
-
A Survey of Continual Reinforcement Learning27 Jun 2025 0 repositories listed
-
Advancements and Challenges in Continual Reinforcement Learning: A Comprehensive Review27 Jun 2025 0 repositories listed
-
Layer Importance for Mathematical Reasoning is Forged in Pre-Training and Invariant after Post-Training27 Jun 2025 0 repositories listed
-
Bridging Offline and Online Reinforcement Learning for LLMs26 Jun 2025 0 repositories listed
-
Explainable AI for Radar Resource Management: Modified LIME in Deep Reinforcement Learning26 Jun 2025 0 repositories listed
-
Quantum Reinforcement Learning Trading Agent for Sector Rotation in the Taiwan Stock Market26 Jun 2025 0 repositories listed
-
Mobile-R1: Towards Interactive Reinforcement Learning for VLM-Based Mobile Agent via Task-Level Rewards25 Jun 2025 0 repositories listed
-
Multi-Objective Reinforcement Learning for Cognitive Radar Resource Management25 Jun 2025 0 repositories listed
-
A Principled Path to Fitted Distributional Evaluation24 Jun 2025 0 repositories listed
-
From Memories to Maps: Mechanisms of In-Context Reinforcement Learning in Transformers24 Jun 2025 0 repositories listed
-
Hierarchical Reinforcement Learning and Value Optimization for Challenging Quadruped Locomotion24 Jun 2025 0 repositories listed
-
Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models24 Jun 2025 0 repositories listed
-
Offline Goal-Conditioned Reinforcement Learning with Projective Quasimetric Planning23 Jun 2025 0 repositories listed
-
Accelerating Residual Reinforcement Learning with Uncertainty Estimation21 Jun 2025 0 repositories listed
-
No Free Lunch: Rethinking Internal Feedback for LLM Reasoning20 Jun 2025 0 repositories listed
-
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning19 Jun 2025 0 repositories listed
-
Multi-Task Lifelong Reinforcement Learning for Wireless Sensor Networks19 Jun 2025 0 repositories listed
-
Multi-Agent Reinforcement Learning for Autonomous Multi-Satellite Earth Observation: A Realistic Case Study18 Jun 2025 0 repositories listed
-
Reinforcement Learning-Based Policy Optimisation For Heterogeneous Radio Access18 Jun 2025 0 repositories listed
-
Steering Your Diffusion Policy with Latent Space Reinforcement Learning18 Jun 2025 0 repositories listed
-
A Comprehensive Survey on Underwater Acoustic Target Positioning and Tracking: Progress, Challenges, and Perspectives17 Jun 2025 0 repositories listed
-
Adaptive Reinforcement Learning for Unobservable Random Delays17 Jun 2025 0 repositories listed
-
HiLight: A Hierarchical Reinforcement Learning Framework with Global Adversarial Guidance for Large-Scale Traffic Signal Control17 Jun 2025 0 repositories listed
-
PeRL: Permutation-Enhanced Reinforcement Learning for Interleaved Vision-Language Reasoning17 Jun 2025 0 repositories listed
-
Situational-Constrained Sequential Resources Allocation via Reinforcement Learning17 Jun 2025 0 repositories listed
-
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models16 Jun 2025 0 repositories listed
-
Discovering Temporal Structure: An Overview of Hierarchical Reinforcement Learning16 Jun 2025 0 repositories listed
-
Dynamic Reinsurance Treaty Bidding via Multi-Agent Reinforcement Learning16 Jun 2025 0 repositories listed
-
Efficient Medical VIE via Reinforcement Learning16 Jun 2025 0 repositories listed
-
Scaling Algorithm Distillation for Continuous Control with Mamba16 Jun 2025 0 repositories listed
-
Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation16 Jun 2025 0 repositories listed
-
Flow-Based Policy for Online Reinforcement Learning15 Jun 2025 0 repositories listed
-
Noise tolerance via reinforcement: Learning a reinforced quantum dynamics14 Jun 2025 0 repositories listed
-
Relative Entropy Regularized Reinforcement Learning for Efficient Encrypted Policy Synthesis14 Jun 2025 0 repositories listed
-
Wasserstein-Barycenter Consensus for Cooperative Multi-Agent Reinforcement Learning14 Jun 2025 0 repositories listed
-
ReVeal: Self-Evolving Code Agents via Iterative Generation-Verification13 Jun 2025 0 repositories listed
-
Efficient Preference-Based Reinforcement Learning: Randomized Exploration Meets Experimental Design11 Jun 2025 0 repositories listed
-
MOORL: A Framework for Integrating Offline-Online Reinforcement Learning11 Jun 2025 0 repositories listed
-
Synergizing Reinforcement Learning and Genetic Algorithms for Neural Combinatorial Optimization11 Jun 2025 0 repositories listed
-
Dynamical System Optimization10 Jun 2025 0 repositories listed
-
Policy-Based Trajectory Clustering in Offline Reinforcement Learning10 Jun 2025 0 repositories listed
-
Predictive reinforcement learning based adaptive PID controller10 Jun 2025 0 repositories listed
-
Semi-gradient DICE for Offline Constrained Reinforcement Learning10 Jun 2025 0 repositories listed
-
TGRPO :Fine-tuning Vision-Language-Action Model via Trajectory-wise Group Relative Policy Optimization10 Jun 2025 0 repositories listed
-
Towards Robust Deep Reinforcement Learning against Environmental State Perturbation10 Jun 2025 0 repositories listed
-
An Intelligent Fault Self-Healing Mechanism for Cloud AI Systems via Integration of Large Language Models and Deep Reinforcement Learning9 Jun 2025 0 repositories listed
-
Decentralizing Multi-Agent Reinforcement Learning with Temporal Causal Information9 Jun 2025 0 repositories listed
-
FairDICE: Fairness-Driven Offline Multi-Objective Reinforcement Learning9 Jun 2025 0 repositories listed
-
ProteinZero: Self-Improving Protein Generation via Online Reinforcement Learning9 Jun 2025 0 repositories listed
-
Reinforcement Learning via Implicit Imitation Guidance9 Jun 2025 0 repositories listed
-
QForce-RL: Quantized FPGA-Optimized Reinforcement Learning Compute Engine8 Jun 2025 0 repositories listed
-
Improving choice model specification using reinforcement learning6 Jun 2025 0 repositories listed
-
Beyond Accuracy: Dissecting Mathematical Reasoning for LLMs Under Reinforcement Learning5 Jun 2025 0 repositories listed
-
Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models5 Jun 2025 0 repositories listed
-
Customizing Speech Recognition Model with Large Language Model Feedback5 Jun 2025 0 repositories listed
-
Discounting and Drug Seeking in Biological Hierarchical Reinforcement Learning5 Jun 2025 0 repositories listed
-
A Lyapunov Drift-Plus-Penalty Method Tailored for Reinforcement Learning with Queue Stability4 Jun 2025 0 repositories listed
-
A Risk-Aware Reinforcement Learning Reward for Financial Trading4 Jun 2025 0 repositories listed
-
Autonomous Vehicle Lateral Control Using Deep Reinforcement Learning with MPC-PID Demonstration4 Jun 2025 0 repositories listed
Syntology lines on 2 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.