Browse State-of-the-Art › Reinforcement Learning (RL) › Papers, page 61
Reinforcement Learning (RL)
Papers archive 2025-07-28
archive papers tagged: 15,113 · with a code link: 4,749 · where Syntology ran a sample: 1,416 (1,186 with a run with no instrument failure, 230 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (1,416 of 15,113 tagged: 1,186 with a run with no instrument failure, 230 where every run was a failure of Syntology's instrument)
Page 61 of 152: papers 6,001 to 6,100 of 15,113, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Pessimism Meets Risk: Risk-Sensitive Offline Reinforcement Learning10 Jul 2024 0 repositories listed
-
Token-Mol 1.0: Tokenized drug design with large language model10 Jul 2024 0 repositories listed
-
Intercepting Unauthorized Aerial Robots in Controlled Airspace Using Reinforcement Learning9 Jul 2024 0 repositories listed
-
An open source Multi-Agent Deep Reinforcement Learning Routing Simulator for satellite networks8 Jul 2024 0 repositories listed
-
Multi-agent Reinforcement Learning-based Network Intrusion Detection System8 Jul 2024 0 repositories listed
-
On Bellman equations for continuous-time policy evaluation I: discretization and approximation8 Jul 2024 0 repositories listed
-
Periodic agent-state based Q-learning for POMDPs8 Jul 2024 0 repositories listed
-
FOSP: Fine-tuning Offline Safe Policy through World Models6 Jul 2024 0 repositories listed
-
Multi-agent Off-policy Actor-Critic Reinforcement Learning for Partially Observable Environments6 Jul 2024 0 repositories listed
-
Autoverse: An Evolvable Game Language for Learning Robust Embodied Agents5 Jul 2024 0 repositories listed
-
Robust Decision Transformer: Tackling Data Corruption in Offline RL via Sequence Modeling5 Jul 2024 0 repositories listed
-
Using Petri Nets as an Integrated Constraint Mechanism for Reinforcement Learning Tasks5 Jul 2024 0 repositories listed
-
Warm-up Free Policy Optimization: Improved Regret in Linear Markov Decision Processes3 Jul 2024 0 repositories listed
-
PWM: Policy Learning with Multi-Task World Models2 Jul 2024 0 repositories listed
-
Reinforcement Learning-driven Data-intensive Workflow Scheduling for Volunteer Edge-Cloud1 Jul 2024 0 repositories listed
-
To Switch or Not to Switch? Balanced Policy Switching in Offline Reinforcement Learning1 Jul 2024 0 repositories listed
-
Benchmarks for Reinforcement Learning with Biased Offline Data and Imperfect Simulators30 Jun 2024 0 repositories listed
-
Safe Reinforcement Learning for Power System Control: A Review30 Jun 2024 0 repositories listed
-
Model-based Offline Reinforcement Learning with Lower Expectile Q-Learning30 Jun 2024 0 repositories listed
-
A Review of Safe Reinforcement Learning Methods for Modern Power Systems29 Jun 2024 0 repositories listed
-
Digital Twin-Assisted Data-Driven Optimization for Reliable Edge Caching in Wireless Networks29 Jun 2024 0 repositories listed
-
Medical Knowledge Integration into Reinforcement Learning Algorithms for Dynamic Treatment Regimes29 Jun 2024 0 repositories listed
-
Beyond Human Preferences: Exploring Reinforcement Learning Trajectory Evaluation and Improvement through LLMs28 Jun 2024 0 repositories listed
-
Decision Transformer for IRS-Assisted Systems with Diffusion-Driven Generative Channels28 Jun 2024 0 repositories listed
-
Optimizing Cyber Defense in Dynamic Active Directories through Reinforcement Learning28 Jun 2024 0 repositories listed
-
Reinforcement Learning for Efficient Design and Control Co-optimisation of Energy Systems28 Jun 2024 0 repositories listed
-
Contrastive Policy Gradient: Aligning LLMs on sequence-level scores in a supervised-friendly fashion27 Jun 2024 0 repositories listed
-
Meta-Gradient Search Control: A Method for Improving the Efficiency of Dyna-style Planning27 Jun 2024 0 repositories listed
-
Decentralized Semantic Traffic Control in AVs Using RL and DQN for Dynamic Roadblocks26 Jun 2024 0 repositories listed
-
Preference Elicitation for Offline Reinforcement Learning26 Jun 2024 0 repositories listed
-
Reinforcement Learning with Intrinsically Motivated Feedback Graph for Lost-sales Inventory Control26 Jun 2024 0 repositories listed
-
EXTRACT: Efficient Policy Learning by Extracting Transferable Robot Skills from Offline Data25 Jun 2024 0 repositories listed
-
Human-Object Interaction from Human-Level Instructions25 Jun 2024 0 repositories listed
-
Leveraging Reinforcement Learning in Red Teaming for Advanced Ransomware Attack Simulations25 Jun 2024 0 repositories listed
-
Privacy Preserving Reinforcement Learning for Population Processes25 Jun 2024 0 repositories listed
-
The State-Action-Reward-State-Action Algorithm in Spatial Prisoner's Dilemma Game25 Jun 2024 0 repositories listed
-
Decentralized RL-Based Data Transmission Scheme for Energy Efficient Harvesting24 Jun 2024 0 repositories listed
-
OCALM: Object-Centric Assessment with Language Models24 Jun 2024 0 repositories listed
-
Tolerance of Reinforcement Learning Controllers against Deviations in Cyber Physical Systems24 Jun 2024 0 repositories listed
-
Diffusion Spectral Representation for Reinforcement Learning23 Jun 2024 0 repositories listed
-
Multistep Criticality Search and Power Shaping in Microreactors with Reinforcement Learning22 Jun 2024 0 repositories listed
-
KalMamba: Towards Efficient Probabilistic State Space Models for RL under Uncertainty21 Jun 2024 0 repositories listed
-
Open Problem: Order Optimal Regret Bounds for Kernel-Based Reinforcement Learning21 Jun 2024 0 repositories listed
-
Equivariant Offline Reinforcement Learning20 Jun 2024 0 repositories listed
-
Optimizing Novelty of Top-k Recommendations using Large Language Models and Reinforcement Learning20 Jun 2024 0 repositories listed
-
Resource Optimization for Tail-Based Control in Wireless Networked Control Systems20 Jun 2024 0 repositories listed
-
Revealing the learning process in reinforcement learning agents through attention-oriented metrics20 Jun 2024 0 repositories listed
-
Rewarding What Matters: Step-by-Step Reinforcement Learning for Task-Oriented Dialogue20 Jun 2024 0 repositories listed
-
Urban-Focused Multi-Task Offline Reinforcement Learning with Contrastive Data Sharing20 Jun 2024 0 repositories listed
-
Learned Graph Rewriting with Equality Saturation: A New Paradigm in Relational Query Rewrite and Beyond19 Jun 2024 0 repositories listed
-
Optimizing Wireless Discontinuous Reception via MAC Signaling Learning19 Jun 2024 0 repositories listed
-
Adaptive Safe Reinforcement Learning-Enabled Optimization of Battery Fast-Charging Protocols18 Jun 2024 0 repositories listed
-
Autonomous navigation of catheters and guidewires in mechanical thrombectomy using inverse reinforcement learning18 Jun 2024 0 repositories listed
-
Order-Optimal Instance-Dependent Bounds for Offline Reinforcement Learning with Preference Feedback18 Jun 2024 0 repositories listed
-
Quantum Compiling with Reinforcement Learning on a Superconducting Processor18 Jun 2024 0 repositories listed
-
Physics-informed Imitative Reinforcement Learning for Real-world Driving18 Jun 2024 0 repositories listed
-
17 Jun 2024 0 repositories listed Syntology 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
Constrained Reinforcement Learning with Average Reward Objective: Model-Based and Model-Free Algorithms17 Jun 2024 0 repositories listed
-
Constructing Ancestral Recombination Graphs through Reinforcement Learning17 Jun 2024 0 repositories listed
-
Linear Bellman Completeness Suffices for Efficient Online Reinforcement Learning with Few Actions17 Jun 2024 0 repositories listed
-
Run Time Assured Reinforcement Learning for Six Degree-of-Freedom Spacecraft Inspection17 Jun 2024 0 repositories listed
-
Design of Interacting Particle Systems for Fast Linear Quadratic RL16 Jun 2024 0 repositories listed
-
Generating and Evolving Reward Functions for Highway Driving with Large Language Models15 Jun 2024 0 repositories listed
-
Finite-Time Analysis of Simultaneous Double Q-learning14 Jun 2024 0 repositories listed
-
ROAR: Reinforcing Original to Augmented Data Ratio Dynamics for Wav2Vec2.0 Based ASR14 Jun 2024 0 repositories listed
-
Unlock the Correlation between Supervised Fine-Tuning and Reinforcement Learning in Training Code Large Language Models14 Jun 2024 0 repositories listed
-
Adaptive Actor-Critic Based Optimal Regulation for Drift-Free Uncertain Nonlinear Systems13 Jun 2024 0 repositories listed
-
CIMRL: Combining IMitation and Reinforcement Learning for Safe Autonomous Driving13 Jun 2024 0 repositories listed
-
Data-driven modeling and supervisory control system optimization for plug-in hybrid electric vehicles13 Jun 2024 0 repositories listed
-
DiffPoGAN: Diffusion Policies with Generative Adversarial Networks for Offline Reinforcement Learning13 Jun 2024 0 repositories listed
-
e-COP : Episodic Constrained Optimization of Policies13 Jun 2024 0 repositories listed
-
SeMOPO: Learning High-quality Model and Policy from Low-quality Offline Visual Datasets13 Jun 2024 0 repositories listed
-
RILe: Reinforced Imitation Learning12 Jun 2024 0 repositories listed
-
Scaling Value Iteration Networks to 5000 Layers for Extreme Long-Term Planning12 Jun 2024 0 repositories listed
-
Toward Enhanced Reinforcement Learning-Based Resource Management via Digital Twin: Opportunities, Applications, and Challenges12 Jun 2024 0 repositories listed
-
CHARME: A chain-based reinforcement learning approach for the minor embedding problem11 Jun 2024 0 repositories listed
-
Enhanced Gene Selection in Single-Cell Genomics: Pre-Filtering Synergy and Reinforced Optimization11 Jun 2024 0 repositories listed
-
Hybrid Reinforcement Learning from Offline Observation Alone11 Jun 2024 0 repositories listed
-
Integrating Domain Knowledge for handling Limited Data in Offline RL11 Jun 2024 0 repositories listed
-
Learning Reward and Policy Jointly from Demonstration and Preference Improves Alignment11 Jun 2024 0 repositories listed
-
Sample Complexity Reduction via Policy Difference Estimation in Tabular Reinforcement Learning11 Jun 2024 0 repositories listed
-
Discovering Multiple Solutions from a Single Task in Offline Reinforcement Learning10 Jun 2024 0 repositories listed
-
Diffusion-based Reinforcement Learning for Dynamic UAV-assisted Vehicle Twins Migration in Vehicular Metaverses8 Jun 2024 0 repositories listed
-
Enhanced Flight Envelope Protection: A Novel Reinforcement Learning Approach8 Jun 2024 0 repositories listed
-
Primitive Agentic First-Order Optimization7 Jun 2024 0 repositories listed
-
Sim-to-Real Transfer of Deep Reinforcement Learning Agents for Online Coverage Path Planning7 Jun 2024 0 repositories listed
-
ATraDiff: Accelerating Online Reinforcement Learning with Imaginary Trajectories6 Jun 2024 0 repositories listed
-
Bootstrapping Expectiles in Reinforcement Learning6 Jun 2024 0 repositories listed
-
Breeding Programs Optimization with Reinforcement Learning6 Jun 2024 0 repositories listed
-
Excluding the Irrelevant: Focusing Reinforcement Learning through Continuous Action Masking6 Jun 2024 0 repositories listed
-
Optimizing Autonomous Driving for Safety: A Human-Centric Approach with LLM-Enhanced RLHF6 Jun 2024 0 repositories listed
-
Proofread: Fixes All Errors with One Tap6 Jun 2024 0 repositories listed
-
Self-Play with Adversarial Critic: Provable and Scalable Offline Alignment for Language Models6 Jun 2024 0 repositories listed
-
DEER: A Delay-Resilient Framework for Reinforcement Learning with Variable Delays5 Jun 2024 0 repositories listed
-
From Tarzan to Tolkien: Controlling the Language Proficiency Level of LLMs for Content Generation5 Jun 2024 0 repositories listed
-
Prompt-based Visual Alignment for Zero-shot Policy Transfer5 Jun 2024 0 repositories listed
-
5 Jun 2024 0 repositories listed Syntology 10 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
UDQL: Bridging The Gap between MSE Loss and The Optimal Value Function in Offline Reinforcement Learning5 Jun 2024 0 repositories listed
-
A Unifying Framework for Action-Conditional Self-Predictive Reinforcement Learning4 Jun 2024 0 repositories listed
Syntology lines on 3 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.