Browse State-of-the-Art › Reinforcement Learning (RL) › Papers, page 48
Reinforcement Learning (RL)
Papers archive 2025-07-28
archive papers tagged: 15,113 · with a code link: 4,749 · where Syntology ran a sample: 1,416 (1,186 with a run with no instrument failure, 230 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (1,416 of 15,113 tagged: 1,186 with a run with no instrument failure, 230 where every run was a failure of Syntology's instrument)
Page 48 of 152: papers 4,701 to 4,800 of 15,113, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
1 Dec 2016 1 repository listed
-
26 Nov 2016 1 repository listed
-
26 Nov 2016 1 repository listed
-
25 Nov 2016 1 repository listed
-
22 Nov 2016 1 repository listed
-
11 Nov 2016 1 repository listed
-
11 Nov 2016 1 repository listed
-
5 Nov 2016 1 repository listed
-
13 Oct 2016 1 repository listed
-
6 Oct 2016 1 repository listed
-
3 Oct 2016 1 repository listed
-
27 Sep 2016 1 repository listed Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
18 Sep 2016 1 repository listed
-
12 Sep 2016 1 repository listed
-
3 Sep 2016 1 repository listed
-
9 Aug 2016 1 repository listed
-
18 Jul 2016 1 repository listed
-
12 Jun 2016 1 repository listed
-
8 Jun 2016 1 repository listed Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
8 Jun 2016 1 repository listed
-
6 Jun 2016 1 repository listed
-
30 May 2016 1 repository listed
-
25 Mar 2016 1 repository listed
-
14 Mar 2016 1 repository listed
-
28 Feb 2016 1 repository listed
-
18 Jan 2016 1 repository listed
-
6 Jan 2016 1 repository listed
-
4 Dec 2015 1 repository listed
-
25 Nov 2015 1 repository listed
-
19 Nov 2015 1 repository listed
-
19 Nov 2015 1 repository listed
-
21 Sep 2015 1 repository listed
-
31 Jul 2015 1 repository listed
-
3 Jul 2015 1 repository listed
-
9 May 2015 1 repository listed
-
4 May 2015 1 repository listed
-
10 Feb 2015 1 repository listed
-
22 Jun 2014 1 repository listed
-
30 Apr 2014 1 repository listed
-
15 Apr 2014 1 repository listed
-
4 Apr 2014 1 repository listed
-
18 Feb 2014 1 repository listed
-
4 Feb 2014 1 repository listed
-
1 Dec 2013 1 repository listed
-
18 Jun 2012 1 repository listed
-
1 Dec 2011 1 repository listed
-
1 Jul 2008 1 repository listed
-
4 Dec 2003 1 repository listed
-
Aligning Humans and Robots via Reinforcement Learning from Implicit Human Feedback17 Jul 2025 0 repositories listed
-
From Novelty to Imitation: Self-Distilled Rewards for Offline Reinforcement Learning17 Jul 2025 0 repositories listed
-
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities17 Jul 2025 0 repositories listed
-
17 Jul 2025 0 repositories listed Syntology 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 12 harvested samples) · 12 pointer-only (licence)
-
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)17 Jul 2025 0 repositories listed
-
VAR-MATH: Probing True Mathematical Reasoning in Large Language Models via Symbolic Multi-Instance Benchmarks17 Jul 2025 0 repositories listed
-
Fly, Fail, Fix: Iterative Game Repair with Reinforcement Learning and Large Multimodal Models16 Jul 2025 0 repositories listed
-
Kevin: Multi-Turn RL for Generating CUDA Kernels16 Jul 2025 0 repositories listed
-
Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training16 Jul 2025 0 repositories listed
-
Illuminating the Three Dogmas of Reinforcement Learning under Evolutionary Light15 Jul 2025 0 repositories listed
-
Local Pairwise Distance Matching for Backpropagation-Free Reinforcement Learning15 Jul 2025 0 repositories listed
-
Real-Time Bayesian Detection of Drift-Evasive GNSS Spoofing in Reinforcement Learning Based UAV Deconfliction15 Jul 2025 0 repositories listed
-
The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs10 Jul 2025 0 repositories listed
-
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model9 Jul 2025 0 repositories listed
-
Video-RTS: Rethinking Reinforcement Learning and Test-Time Scaling for Efficient and Enhanced Video Reasoning9 Jul 2025 0 repositories listed
-
CogniSQL-R1-Zero: Lightweight Reinforced Reasoning for Efficient SQL Generation8 Jul 2025 0 repositories listed
-
Detecting and Mitigating Reward Hacking in Reinforcement Learning Systems: A Comprehensive Empirical Study8 Jul 2025 0 repositories listed
-
FEVO: Financial Knowledge Expansion and Reasoning Evolution for Large Language Models8 Jul 2025 0 repositories listed
-
Robust Bandwidth Estimation for Real-Time Communication with Offline Reinforcement Learning8 Jul 2025 0 repositories listed
-
Safe Domain Randomization via Uncertainty-Aware Out-of-Distribution Detection and Policy Adaptation8 Jul 2025 0 repositories listed
-
2048: Reinforcement Learning in a Delayed Reward Environment7 Jul 2025 0 repositories listed
-
Open Vision Reasoner: Transferring Linguistic Cognitive Behavior for Visual Reasoning7 Jul 2025 0 repositories listed
-
Listener-Rewarded Thinking in VLMs for Image Preferences28 Jun 2025 0 repositories listed
-
A Survey of Continual Reinforcement Learning27 Jun 2025 0 repositories listed
-
Advancements and Challenges in Continual Reinforcement Learning: A Comprehensive Review27 Jun 2025 0 repositories listed
-
APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization26 Jun 2025 0 repositories listed
-
Curriculum-Guided Antifragile Reinforcement Learning for Secure UAV Deconfliction under Observation-Space Attacks26 Jun 2025 0 repositories listed
-
26 Jun 2025 0 repositories listed Syntology 3 ran (of which 2 constructed an object rather than computing a result; 3 with no instrument failure: 1 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Optimising 4th-Order Runge-Kutta Methods: A Dynamic Heuristic Approach for Efficiency and Low Storage26 Jun 2025 0 repositories listed
-
RL-Selector: Reinforcement Learning-Guided Data Selection via Redundancy Assessment26 Jun 2025 0 repositories listed
-
Robust Policy Switching for Antifragile Reinforcement Learning for UAV Deconfliction in Adversarial Environments26 Jun 2025 0 repositories listed
-
Strict Subgoal Execution: Reliable Long-Horizon Planning in Hierarchical Reinforcement Learning26 Jun 2025 0 repositories listed
-
Asymmetric REINFORCE for off-Policy Reinforcement Learning: Balancing positive and negative rewards25 Jun 2025 0 repositories listed
-
A Comparative Analysis of Reinforcement Learning and Conventional Deep Learning Approaches for Bearing Fault Diagnosis24 Jun 2025 0 repositories listed
-
Causal-Aware Intelligent QoE Optimization for VR Interaction with Adaptive Keyframe Extraction24 Jun 2025 0 repositories listed
-
Hierarchical Reinforcement Learning and Value Optimization for Challenging Quadruped Locomotion24 Jun 2025 0 repositories listed
-
AdapThink: Adaptive Thinking Preferences for Reasoning Language Model23 Jun 2025 0 repositories listed
-
Robots and Children that Learn Together : Improving Knowledge Retention by Teaching Peer-Like Interactive Robots23 Jun 2025 0 repositories listed
-
Accelerating Residual Reinforcement Learning with Uncertainty Estimation21 Jun 2025 0 repositories listed
-
Leveling the Playing Field: Carefully Comparing Classical and Learned Controllers for Quadrotor Trajectory Tracking21 Jun 2025 0 repositories listed
-
Learning Dexterous Object Handover20 Jun 2025 0 repositories listed
-
Dual-Objective Reinforcement Learning with Novel Hamilton-Jacobi-Bellman Formulations19 Jun 2025 0 repositories listed
-
From General to Targeted Rewards: Surpassing GPT-4 in Open-Ended Long-Context Generation19 Jun 2025 0 repositories listed
-
Multi-Task Lifelong Reinforcement Learning for Wireless Sensor Networks19 Jun 2025 0 repositories listed
-
VRAIL: Vectorized Reward-based Attribution for Interpretable Learning19 Jun 2025 0 repositories listed
-
Make Your AUV Adaptive: An Environment-Aware Reinforcement Learning Framework For Underwater Tasks18 Jun 2025 0 repositories listed
-
Multi-Agent Reinforcement Learning for Autonomous Multi-Satellite Earth Observation: A Realistic Case Study18 Jun 2025 0 repositories listed
-
Reinforcement Learning-Based Policy Optimisation For Heterogeneous Radio Access18 Jun 2025 0 repositories listed
-
Steering Your Diffusion Policy with Latent Space Reinforcement Learning18 Jun 2025 0 repositories listed
-
Adaptive Reinforcement Learning for Unobservable Random Delays17 Jun 2025 0 repositories listed
-
HiLight: A Hierarchical Reinforcement Learning Framework with Global Adversarial Guidance for Large-Scale Traffic Signal Control17 Jun 2025 0 repositories listed
Syntology lines on 5 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.