Browse State-of-the-Art › Reinforcement Learning (RL) › Papers, page 58
Reinforcement Learning (RL)
Papers archive 2025-07-28
archive papers tagged: 15,113 · with a code link: 4,749 · where Syntology ran a sample: 1,416 (1,186 with a run with no instrument failure, 230 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (1,416 of 15,113 tagged: 1,186 with a run with no instrument failure, 230 where every run was a failure of Syntology's instrument)
Page 58 of 152: papers 5,701 to 5,800 of 15,113, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Learn 2 Rage: Experiencing The Emotional Roller Coaster That Is Reinforcement Learning24 Oct 2024 0 repositories listed
-
PointPatchRL -- Masked Reconstruction Improves Reinforcement Learning on Point Clouds24 Oct 2024 0 repositories listed
-
SAMG: State-Action-Aware Offline-to-Online Reinforcement Learning with Offline Model Guidance24 Oct 2024 0 repositories listed
-
The Hive Mind is a Single Reinforcement Learning Agent23 Oct 2024 0 repositories listed
-
Dynamic Spectrum Access for Ambient Backscatter Communication-assisted D2D Systems with Quantum Reinforcement Learning23 Oct 2024 0 repositories listed
-
Optimizing Load Scheduling in Power Grids Using Reinforcement Learning and Markov Decision Processes23 Oct 2024 0 repositories listed
-
Primal-Dual Spectral Representation for Off-policy Evaluation23 Oct 2024 0 repositories listed
-
Process Supervision-Guided Policy Optimization for Code Generation23 Oct 2024 0 repositories listed
-
Benchmarking Smoothness and Reducing High-Frequency Oscillations in Continuous Control Policies22 Oct 2024 0 repositories listed
-
DROP: Distributional and Regular Optimism and Pessimism for Reinforcement Learning22 Oct 2024 0 repositories listed
-
DyPNIPP: Predicting Environment Dynamics for RL-based Robust Informative Path Planning22 Oct 2024 0 repositories listed
-
Episodic Future Thinking Mechanism for Multi-agent Reinforcement Learning22 Oct 2024 0 repositories listed
-
Meta Stackelberg Game: Robust Federated Learning against Adaptive and Mixed Poisoning Attacks22 Oct 2024 0 repositories listed
-
Multi-Modal Transformer and Reinforcement Learning-based Beam Management22 Oct 2024 0 repositories listed
-
Curriculum Reinforcement Learning for Complex Reward Functions22 Oct 2024 0 repositories listed
-
Survival of the Fittest: Evolutionary Adaptation of Policies for Environmental Shifts22 Oct 2024 0 repositories listed
-
Offline reinforcement learning for job-shop scheduling problems21 Oct 2024 0 repositories listed
-
Training Language Models to Critique With Multi-agent Feedback20 Oct 2024 0 repositories listed
-
A Novel Reinforcement Learning Model for Post-Incident Malware Investigations19 Oct 2024 0 repositories listed
-
Action abstractions for amortized sampling19 Oct 2024 0 repositories listed
-
Augmented Lagrangian-Based Safe Reinforcement Learning Approach for Distribution System Volt/VAR Control19 Oct 2024 0 repositories listed
-
MENTOR: Mixture-of-Experts Network with Task-Oriented Perturbation for Visual Reinforcement Learning19 Oct 2024 0 repositories listed
-
A Large Language Model-Driven Reward Design Framework via Dynamic Feedback for Reinforcement Learning18 Oct 2024 0 repositories listed
-
Harnessing Causality in Reinforcement Learning With Bagged Decision Times18 Oct 2024 0 repositories listed
-
Interpretable end-to-end Neurosymbolic Reinforcement Learning agents18 Oct 2024 0 repositories listed
-
Reinforcement Learning in Non-Markov Market-Making18 Oct 2024 0 repositories listed
-
Coordinated Dispatch of Energy Storage Systems in the Active Distribution Network: A Complementary Reinforcement Learning and Optimization Approach17 Oct 2024 0 repositories listed
-
Guided Reinforcement Learning for Robust Multi-Contact Loco-Manipulation17 Oct 2024 0 repositories listed
-
Integrating Large Language Models and Reinforcement Learning for Non-Linear Reasoning17 Oct 2024 0 repositories listed
-
MarineFormer: A Spatio-Temporal Attention Model for USV Navigation in Dynamic Marine Environments17 Oct 2024 0 repositories listed
-
Truncating Trajectories in Monte Carlo Policy Evaluation: an Adaptive Approach17 Oct 2024 0 repositories listed
-
Augmented Intelligence in Smart Intersections: Local Digital Twins-Assisted Hybrid Autonomous Driving16 Oct 2024 0 repositories listed
-
Dynamic Learning Rate for Deep Reinforcement Learning: A Bandit Approach16 Oct 2024 0 repositories listed
-
EdgeRL: Reinforcement Learning-driven Deep Learning Model Inference Optimization at Edge16 Oct 2024 0 repositories listed
-
Neural-based Control for CubeSat Docking Maneuvers16 Oct 2024 0 repositories listed
-
Off-dynamics Conditional Diffusion Planners16 Oct 2024 0 repositories listed
-
Reinforcement Learning with LTL and ω-Regular Objectives via Optimality-Preserving Translation to Average Rewards16 Oct 2024 0 repositories listed
-
Robust RL with LLM-Driven Data Synthesis and Policy Adaptation for Autonomous Driving16 Oct 2024 0 repositories listed
-
SAC-GLAM: Improving Online RL for LLM agents with Soft Actor-Critic and Hindsight Relabeling16 Oct 2024 0 repositories listed
-
When to Trust Your Data: Enhancing Dyna-Style Model-Based Reinforcement Learning With Data Filter16 Oct 2024 0 repositories listed
-
Diffusion-Based Offline RL for Improved Decision-Making in Augmented ARC Task15 Oct 2024 0 repositories listed
-
ILAEDA: An Imitation Learning Based Approach for Automatic Exploratory Data Analysis15 Oct 2024 0 repositories listed
-
Multi-Objective-Optimization Multi-AUV Assisted Data Collection Framework for IoUT Based on Offline Reinforcement Learning15 Oct 2024 0 repositories listed
-
Multi-objective Reinforcement Learning: A Tool for Pluralistic Alignment15 Oct 2024 0 repositories listed
-
Reinforcement Learning Based Bidding Framework with High-dimensional Bids in Power Markets15 Oct 2024 0 repositories listed
-
Burning RED: Unlocking Subtask-Driven Reinforcement Learning and Risk-Awareness in Average-Reward Markov Decision Processes14 Oct 2024 0 repositories listed
-
DR-MPC: Deep Residual Model Predictive Control for Real-world Social Navigation14 Oct 2024 0 repositories listed
-
Large Language Model-Enhanced Reinforcement Learning for Generic Bus Holding Control Strategies14 Oct 2024 0 repositories listed
-
Asymptotic Analysis of Sample-averaged Q-learning14 Oct 2024 0 repositories listed
-
Generalization of Compositional Tasks with Logical Specification via Implicit Planning13 Oct 2024 0 repositories listed
-
Integrating Reinforcement Learning and Large Language Models for Crop Production Process Management Optimization and Control through A New Knowledge-Based Deep Learning Paradigm13 Oct 2024 0 repositories listed
-
Meta-Reinforcement Learning with Universal Policy Adaptation: Provable Near-Optimality under All-task Optimum Comparator13 Oct 2024 0 repositories listed
-
Transformers as Game Players: Provable In-context Game-playing Capabilities of Pre-trained Models13 Oct 2024 0 repositories listed
-
ActSafe: Active Exploration with Safety Constraints for Reinforcement Learning12 Oct 2024 0 repositories listed
-
Reinforcement Learning in Hyperbolic Spaces: Models and Experiments12 Oct 2024 0 repositories listed
-
Can we hop in general? A discussion of benchmark selection and design using the Hopper environment11 Oct 2024 0 repositories listed
-
MAD-TD: Model-Augmented Data stabilizes High Update Ratio RL11 Oct 2024 0 repositories listed
-
Physical Simulation for Multi-agent Multi-machine Tending11 Oct 2024 0 repositories listed
-
SOLD: Slot Object-Centric Latent Dynamics Models for Relational Manipulation Learning from Pixels11 Oct 2024 0 repositories listed
-
Words as Beacons: Guiding RL Agents with High-Level Language Prompts11 Oct 2024 0 repositories listed
-
Avoiding mode collapse in diffusion models fine-tuned with reinforcement learning10 Oct 2024 0 repositories listed
-
Efficient Reinforcement Learning with Large Language Model Priors10 Oct 2024 0 repositories listed
-
Masked Generative Priors Improve World Models Sequence Modelling Capabilities10 Oct 2024 0 repositories listed
-
Offline Hierarchical Reinforcement Learning via Inverse Optimization10 Oct 2024 0 repositories listed
-
Offline Inverse Constrained Reinforcement Learning for Safe-Critical Decision Making in Healthcare10 Oct 2024 0 repositories listed
-
Probabilistic Satisfaction of Temporal Logic Constraints in Reinforcement Learning via Adaptive Policy-Switching10 Oct 2024 0 repositories listed
-
Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning10 Oct 2024 0 repositories listed
-
VerifierQ: Enhancing LLM Test Time Compute with Q-Learning-based Verifiers10 Oct 2024 0 repositories listed
-
A Safety Modulator Actor-Critic Method in Model-Free Safe Reinforcement Learning and Application in UAV Hovering9 Oct 2024 0 repositories listed
-
Flipping-based Policy for Chance-Constrained Markov Decision Processes9 Oct 2024 0 repositories listed
-
MotionRL: Align Text-to-Motion Generation to Human Preferences with Multi-Reward Reinforcement Learning9 Oct 2024 0 repositories listed
-
Q-WSL: Optimizing Goal-Conditioned RL with Weighted Supervised Learning via Dynamic Programming9 Oct 2024 0 repositories listed
-
Zero-Shot Generalization of Vision-Based RL Without Data Augmentation9 Oct 2024 0 repositories listed
-
On the Modeling Capabilities of Large Language Models for Sequential Decision Making8 Oct 2024 0 repositories listed
-
Reinforcement Learning From Imperfect Corrective Actions And Proxy Rewards8 Oct 2024 0 repositories listed
-
Solving Multi-Goal Robotic Tasks with Decision Transformer8 Oct 2024 0 repositories listed
-
Solving robust MDPs as a sequence of static RL problems8 Oct 2024 0 repositories listed
-
AlphaRouter: Quantum Circuit Routing with Reinforcement Learning and Tree Search7 Oct 2024 0 repositories listed
-
Towards Measuring Goal-Directedness in AI Systems7 Oct 2024 0 repositories listed
-
Towards using Reinforcement Learning for Scaling and Data Replication in Cloud Systems7 Oct 2024 0 repositories listed
-
A Reinforcement Learning Engine with Reduced Action and State Space for Scalable Cyber-Physical Optimal Response6 Oct 2024 0 repositories listed
-
AdaMemento: Adaptive Memory-Assisted Policy Optimization for Reinforcement Learning6 Oct 2024 0 repositories listed
-
Data-driven Under Frequency Load Shedding Using Reinforcement Learning6 Oct 2024 0 repositories listed
-
DeepLTL: Learning to Efficiently Satisfy Complex LTL Specifications for Multi-Task RL6 Oct 2024 0 repositories listed
-
4 Oct 2024 0 repositories listed Syntology 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Beyond Expected Returns: A Policy Gradient Algorithm for Cumulative Prospect Theoretic Reinforcement Learning3 Oct 2024 0 repositories listed
-
Cross-Embodiment Dexterous Grasping with Reinforcement Learning3 Oct 2024 0 repositories listed
-
Dual Active Learning for Reinforcement Learning from Human Feedback3 Oct 2024 0 repositories listed
-
Efficient Residual Learning with Mixture-of-Experts for Universal Dexterous Grasping3 Oct 2024 0 repositories listed
-
End-to-end Driving in High-Interaction Traffic Scenarios with Reinforcement Learning3 Oct 2024 0 repositories listed
-
Learning Emergence of Interaction Patterns across Independent RL Agents in Multi-Agent Environments3 Oct 2024 0 repositories listed
-
Solving Reach-Avoid-Stay Problems Using Deep Deterministic Policy Gradients3 Oct 2024 0 repositories listed
-
Absolute State-wise Constrained Policy Optimization: High-Probability State-wise Constraints Satisfaction2 Oct 2024 0 repositories listed
-
Bellman Diffusion: Generative Modeling as Learning a Linear Operator in the Distribution Space2 Oct 2024 0 repositories listed
-
ComaDICE: Offline Cooperative Multi-Agent Reinforcement Learning with Stationary Distribution Shift Regularization2 Oct 2024 0 repositories listed
-
Don't flatten, tokenize! Unlocking the key to SoftMoE's efficacy in deep RL2 Oct 2024 0 repositories listed
-
LLM-Augmented Symbolic Reinforcement Learning with Landmark-Based Task Decomposition2 Oct 2024 0 repositories listed
-
PreND: Enhancing Intrinsic Motivation in Reinforcement Learning through Pre-trained Network Distillation2 Oct 2024 0 repositories listed
-
The Smart Buildings Control Suite: A Diverse Open Source Benchmark to Evaluate and Scale HVAC Control Policies for Sustainability2 Oct 2024 0 repositories listed
Syntology lines on 2 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.