Browse State-of-the-Art › Reinforcement Learning (RL) › Papers, page 52
Reinforcement Learning (RL)
Papers archive 2025-07-28
archive papers tagged: 15,113 · with a code link: 4,749 · where Syntology ran a sample: 1,416 (1,186 with a run with no instrument failure, 230 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (1,416 of 15,113 tagged: 1,186 with a run with no instrument failure, 230 where every run was a failure of Syntology's instrument)
Page 52 of 152: papers 5,101 to 5,200 of 15,113, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Hybrid Reinforcement Learning and Model Predictive Control for Adaptive Control of Hydrogen-Diesel Dual-Fuel Combustion23 Apr 2025 0 repositories listed
-
Monte Carlo Planning with Large Language Model for Text-Based Game Agents23 Apr 2025 0 repositories listed
-
Natural Policy Gradient for Average Reward Non-Stationary RL23 Apr 2025 0 repositories listed
-
Offline Robotic World Model: Learning Robotic Policies without a Physics Simulator23 Apr 2025 0 repositories listed
-
Reinforcement learning framework for the mechanical design of microelectronic components under multiphysics constraints23 Apr 2025 0 repositories listed
-
Insights from Verification: Training a Verilog Generation LLM with Reinforcement Learning with Testbench Feedback22 Apr 2025 0 repositories listed
-
Policy-Based Radiative Transfer: Solving the 2-Level Atom Non-LTE Problem using Soft Actor-Critic Reinforcement Learning22 Apr 2025 0 repositories listed
-
Real-Time Optimal Design of Experiment for Parameter Identification of Li-Ion Cell Electrochemical Model22 Apr 2025 0 repositories listed
-
SARI: Structured Audio Reasoning via Curriculum-Guided Reinforcement Learning22 Apr 2025 0 repositories listed
-
SLiM-Gym: Reinforcement Learning for Population Genetics22 Apr 2025 0 repositories listed
-
StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation22 Apr 2025 0 repositories listed
-
Dynamic Contrastive Skill Learning with State-Transition Based Skill Clustering and Dynamic Length Adjustment21 Apr 2025 0 repositories listed
-
LAPP: Large Language Model Feedback for Preference-Driven Reinforcement Learning21 Apr 2025 0 repositories listed
-
OTC: Optimal Tool Calls via Reinforcement Learning21 Apr 2025 0 repositories listed
-
Think2SQL: Reinforce LLM Reasoning Capabilities for Text2SQL21 Apr 2025 0 repositories listed
-
Relation-R1: Cognitive Chain-of-Thought Guided Reinforcement Learning for Unified Relational Comprehension20 Apr 2025 0 repositories listed
-
Improving RL Exploration for LLM Reasoning through Retrospective Replay19 Apr 2025 0 repositories listed
-
Mixed-Precision Conjugate Gradient Solvers with RL-Driven Precision Tuning19 Apr 2025 0 repositories listed
-
Quantum-Enhanced Reinforcement Learning for Power Grid Security Assessment19 Apr 2025 0 repositories listed
-
Unlearning Works Better Than You Think: Local Reinforcement-Based Selection of Auxiliary Objectives19 Apr 2025 0 repositories listed
-
Improving Generalization in Intent Detection: GRPO with Reward-Based Curriculum Sampling18 Apr 2025 0 repositories listed
-
Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning18 Apr 2025 0 repositories listed
-
SwitchMT: An Adaptive Context Switching Methodology for Scalable Multi-Task Learning in Intelligent Autonomous Agents18 Apr 2025 0 repositories listed
-
Crossing the Human-Robot Embodiment Gap with Sim-to-Real RL using One Human Demonstration17 Apr 2025 0 repositories listed
-
Evolutionary Policy Optimization17 Apr 2025 0 repositories listed
-
LLMs Meet Finance: Fine-Tuning Foundation Models for the Open FinLLM Leaderboard17 Apr 2025 0 repositories listed
-
RL-PINNs: Reinforcement Learning-Driven Adaptive Sampling for Efficient Training of PINNs17 Apr 2025 0 repositories listed
-
TraCeS: Trajectory Based Credit Assignment From Sparse Safety Feedback17 Apr 2025 0 repositories listed
-
d1: Scaling Reasoning in Diffusion Large Language Models via Reinforcement Learning16 Apr 2025 0 repositories listed
-
Evolutionary Reinforcement Learning for Interpretable Decision-Making in Supply Chain Management16 Apr 2025 0 repositories listed
-
pix2pockets: Shot Suggestions in 8-Ball Pool from a Single Image in the Wild16 Apr 2025 0 repositories listed
-
16 Apr 2025 0 repositories listed Syntology 10 ran (of which 9 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified (of 14 harvested samples) · 14 pointer-only (licence)
-
Achieving Tighter Finite-Time Rates for Heterogeneous Federated Stochastic Approximation under Markovian Sampling15 Apr 2025 0 repositories listed
-
Hallucination-Aware Generative Pretrained Transformer for Cooperative Aerial Mobility Control15 Apr 2025 0 repositories listed
-
Next-Future: Sample-Efficient Policy Learning for Robotic-Arm Tasks15 Apr 2025 0 repositories listed
-
Position Paper: Rethinking Privacy in RL for Sequential Decision-making in the Age of LLMs15 Apr 2025 0 repositories listed
-
Revealing Covert Attention by Analyzing Human and Reinforcement Learning Agent Gameplay15 Apr 2025 0 repositories listed
-
ReZero: Enhancing LLM search ability by trying one-more-time15 Apr 2025 0 repositories listed
-
Adaptive Insurance Reserving with CVaR-Constrained Reinforcement Learning under Macroeconomic Regimes13 Apr 2025 0 repositories listed
-
CheatAgent: Attacking LLM-Empowered Recommender Systems via LLM Agent13 Apr 2025 0 repositories listed
-
Efficient Implementation of Reinforcement Learning over Homomorphic Encryption12 Apr 2025 0 repositories listed
-
Towards More Efficient, Robust, Instance-adaptive, and Generalizable Sequential Decision making12 Apr 2025 0 repositories listed
-
Towards Optimal Differentially Private Regret Bounds in Linear MDPs12 Apr 2025 0 repositories listed
-
Deep Distributional Learning with Non-crossing Quantile Network11 Apr 2025 0 repositories listed
-
Spectral Normalization for Lipschitz-Constrained Policies on Learning Humanoid Locomotion11 Apr 2025 0 repositories listed
-
Boosting Universal LLM Reward Design through the Heuristic Reward Observation Space Evolution10 Apr 2025 0 repositories listed
-
Deep Reinforcement Learning for Day-to-day Dynamic Tolling in Tradable Credit Schemes10 Apr 2025 0 repositories listed
-
Fast Adaptation with Behavioral Foundation Models10 Apr 2025 0 repositories listed
-
Genetic Programming with Reinforcement Learning Trained Transformer for Real-World Dynamic Scheduling Problems10 Apr 2025 0 repositories listed
-
RL-based Control of UAS Subject to Significant Disturbance10 Apr 2025 0 repositories listed
-
Better Decisions through the Right Causal World Model9 Apr 2025 0 repositories listed
-
Smart Exploration in Reinforcement Learning using Bounded Uncertainty Models8 Apr 2025 0 repositories listed
-
Stratified Expert Cloning with Adaptive Selection for User Retention in Large-Scale Recommender Systems8 Apr 2025 0 repositories listed
-
TW-CRL: Time-Weighted Contrastive Reward Learning for Efficient Inverse Reinforcement Learning8 Apr 2025 0 repositories listed
-
xMTF: A Formula-Free Model for Reinforcement-Learning-Based Multi-Task Fusion in Recommender Systems8 Apr 2025 0 repositories listed
-
Algorithm Discovery With LLMs: Evolutionary Search Meets Reinforcement Learning7 Apr 2025 0 repositories listed
-
Physics-informed Modularized Neural Network for Advanced Building Control by Deep Reinforcement Learning7 Apr 2025 0 repositories listed
-
The Role of Environment Access in Agnostic Reinforcement Learning7 Apr 2025 0 repositories listed
-
Impact of Price Inflation on Algorithmic Collusion Through Reinforcement Learning Agents5 Apr 2025 0 repositories listed
-
OrbitZoo: Multi-Agent Reinforcement Learning Environment for Orbital Dynamics5 Apr 2025 0 repositories listed
-
Algorithmic Prompt Generation for Diverse Human-like Teaming and Communication with Large Language Models4 Apr 2025 0 repositories listed
-
Decision SpikeFormer: Spike-Driven Transformer for Decision Making4 Apr 2025 0 repositories listed
-
Dexterous Manipulation through Imitation Learning: A Survey4 Apr 2025 0 repositories listed
-
Enhanced Penalty-based Bidirectional Reinforcement Learning Algorithms4 Apr 2025 0 repositories listed
-
Improving Mixed-Criticality Scheduling with Reinforcement Learning4 Apr 2025 0 repositories listed
-
Learning Dual-Arm Coordination for Grasping Large Flat Objects4 Apr 2025 0 repositories listed
-
Offline and Distributional Reinforcement Learning for Wireless Communications4 Apr 2025 0 repositories listed
-
Adapting World Models with Latent-State Dynamics Residuals3 Apr 2025 0 repositories listed
-
Inference-Time Scaling for Generalist Reward Modeling3 Apr 2025 0 repositories listed
-
Integrating Human Knowledge Through Action Masking in Reinforcement Learning for Operations Research3 Apr 2025 0 repositories listed
-
Reinforcement Learning for Solving the Pricing Problem in Column Generation: Applications to Vehicle Routing3 Apr 2025 0 repositories listed
-
De Novo Molecular Design Enabled by Direct Preference Optimization and Curriculum Learning2 Apr 2025 0 repositories listed
-
Probabilistic Curriculum Learning for Goal-Based Reinforcement Learning2 Apr 2025 0 repositories listed
-
Grounding Multimodal LLMs to Embodied Agents that Ask for Help with Reinforcement Learning1 Apr 2025 0 repositories listed
-
How Difficulty-Aware Staged Reinforcement Learning Enhances LLMs' Reasoning Capabilities: A Preliminary Experimental Study1 Apr 2025 0 repositories listed
-
A Survey of Reinforcement Learning-Based Motion Planning for Autonomous Driving: Lessons Learned from a Driving Task Perspective31 Mar 2025 0 repositories listed
-
Accelerating High-Efficiency Organic Photovoltaic Discovery via Pretrained Graph Neural Networks and Generative Reinforcement Learning31 Mar 2025 0 repositories listed
-
Fair Dynamic Spectrum Access via Fully Decentralized Multi-Agent Reinforcement Learning31 Mar 2025 0 repositories listed
-
HACTS: a Human-As-Copilot Teleoperation System for Robot Learning31 Mar 2025 0 repositories listed
-
JudgeLRM: Large Reasoning Models as a Judge31 Mar 2025 0 repositories listed
-
Noise-based reward-modulated learning31 Mar 2025 0 repositories listed
-
Nuclear Microreactor Control with Deep Reinforcement Learning31 Mar 2025 0 repositories listed
-
Reinforcement Learning for Safe Autonomous Two Device Navigation of Cerebral Vessels in Mechanical Thrombectomy31 Mar 2025 0 repositories listed
-
A Systematic Decade Review of Trip Route Planning with Travel Time Estimation based on User Preferences and Behavior30 Mar 2025 0 repositories listed
-
Advanced Deep Learning and Large Language Models: Comprehensive Insights for Cancer Detection30 Mar 2025 0 repositories listed
-
Reinforcement Learning for Active Matter30 Mar 2025 0 repositories listed
-
Multi-Agent Reinforcement Learning for Graph Discovery in D2D-Enabled Federated Learning29 Mar 2025 0 repositories listed
-
Reasoning-SQL: Reinforcement Learning with SQL Tailored Partial Rewards for Reasoning-Enhanced Text-to-SQL29 Mar 2025 0 repositories listed
-
RL2Grid: Benchmarking Reinforcement Learning in Power Grid Operations29 Mar 2025 0 repositories listed
-
Entropy-guided sequence weighting for efficient exploration in RL-based LLM fine-tuning28 Mar 2025 0 repositories listed
-
FLAM: Foundation Model-Based Body Stabilization for Humanoid Locomotion and Manipulation28 Mar 2025 0 repositories listed
-
Reinforcement Learning for Machine Learning Model Deployment: Evaluating Multi-Armed Bandits in ML Ops Environments28 Mar 2025 0 repositories listed
-
Bresa: Bio-inspired Reflexive Safe Reinforcement Learning for Contact-Rich Robotic Tasks27 Mar 2025 0 repositories listed
-
Harmonia: A Multi-Agent Reinforcement Learning Approach to Data Placement and Migration in Hybrid Storage Systems26 Mar 2025 0 repositories listed
-
Learning Adaptive Dexterous Grasping from Single Demonstrations26 Mar 2025 0 repositories listed
-
Model-Based Offline Reinforcement Learning with Adversarial Data Augmentation26 Mar 2025 0 repositories listed
-
Offline Reinforcement Learning with Discrete Diffusion Skills26 Mar 2025 0 repositories listed
-
Reasoning Beyond Limits: Advances and Open Problems for LLMs26 Mar 2025 0 repositories listed
-
Synthesizing world models for bilevel planning26 Mar 2025 0 repositories listed
-
TAR: Teacher-Aligned Representations via Contrastive Learning for Quadrupedal Locomotion26 Mar 2025 0 repositories listed
Syntology lines on 1 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.