Browse State-of-the-Art › reinforcement-learning › Papers, page 53
reinforcement-learning
Papers archive 2025-07-28
archive papers tagged: 13,427 · with a code link: 4,119 · where Syntology ran a sample: 1,165 (973 with a run with no instrument failure, 192 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (1,165 of 13,427 tagged: 973 with a run with no instrument failure, 192 where every run was a failure of Syntology's instrument)
Page 53 of 135: papers 5,201 to 5,300 of 13,427, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Contractual Reinforcement Learning: Pulling Arms with Invisible Hands1 Jul 2024 0 repositories listed
-
Deep Reinforcement Learning for Adverse Garage Scenario Generation1 Jul 2024 0 repositories listed
-
Normalization and effective learning rates in reinforcement learning1 Jul 2024 0 repositories listed
-
Benchmarks for Reinforcement Learning with Biased Offline Data and Imperfect Simulators30 Jun 2024 0 repositories listed
-
Disentangled Representations for Causal Cognition30 Jun 2024 0 repositories listed
-
Safe Reinforcement Learning for Power System Control: A Review30 Jun 2024 0 repositories listed
-
Model-based Offline Reinforcement Learning with Lower Expectile Q-Learning30 Jun 2024 0 repositories listed
-
Medical Knowledge Integration into Reinforcement Learning Algorithms for Dynamic Treatment Regimes29 Jun 2024 0 repositories listed
-
Beyond Human Preferences: Exploring Reinforcement Learning Trajectory Evaluation and Improvement through LLMs28 Jun 2024 0 repositories listed
-
Optimizing Cyber Defense in Dynamic Active Directories through Reinforcement Learning28 Jun 2024 0 repositories listed
-
AI Alignment through Reinforcement Learning from Human Feedback? Contradictions and Limitations26 Jun 2024 0 repositories listed
-
Mental Modeling of Reinforcement Learning Agents by Language Models26 Jun 2024 0 repositories listed
-
Preference Elicitation for Offline Reinforcement Learning26 Jun 2024 0 repositories listed
-
Reinforcement Learning with Intrinsically Motivated Feedback Graph for Lost-sales Inventory Control26 Jun 2024 0 repositories listed
-
Privacy Preserving Reinforcement Learning for Population Processes25 Jun 2024 0 repositories listed
-
Model-Free Robust Reinforcement Learning with Sample Complexity Analysis24 Jun 2024 0 repositories listed
-
Diffusion Spectral Representation for Reinforcement Learning23 Jun 2024 0 repositories listed
-
Position: Benchmarking is Limited in Reinforcement Learning Research23 Jun 2024 0 repositories listed
-
Understanding and Diagnosing Deep Reinforcement Learning23 Jun 2024 0 repositories listed
-
Distributionally Robust Constrained Reinforcement Learning under Strong Duality22 Jun 2024 0 repositories listed
-
An Idiosyncrasy of Time-discretization in Reinforcement Learning21 Jun 2024 0 repositories listed
-
KnobTree: Intelligent Database Parameter Configuration via Explainable Reinforcement Learning21 Jun 2024 0 repositories listed
-
Robust Reinforcement Learning from Corrupted Human Feedback21 Jun 2024 0 repositories listed
-
What Teaches Robots to Walk, Teaches Them to Trade too -- Regime Adaptive Execution using Informed Data and LLMs20 Jun 2024 0 repositories listed
-
A General Control-Theoretic Approach for Reinforcement Learning: Theory and Algorithms20 Jun 2024 0 repositories listed
-
Bayesian Inverse Reinforcement Learning for Non-Markovian Rewards20 Jun 2024 0 repositories listed
-
Constrained Meta Agnostic Reinforcement Learning20 Jun 2024 0 repositories listed
-
Equivariant Offline Reinforcement Learning20 Jun 2024 0 repositories listed
-
Tractable Equilibrium Computation in Markov Games through Risk Aversion20 Jun 2024 0 repositories listed
-
Urban-Focused Multi-Task Offline Reinforcement Learning with Contrastive Data Sharing20 Jun 2024 0 repositories listed
-
Learned Graph Rewriting with Equality Saturation: A New Paradigm in Relational Query Rewrite and Beyond19 Jun 2024 0 repositories listed
-
Infinite-Horizon Reinforcement Learning with Multinomial Logistic Function Approximation19 Jun 2024 0 repositories listed
-
Trapezoidal Gradient Descent for Effective Reinforcement Learning in Spiking Networks19 Jun 2024 0 repositories listed
-
Quantum Compiling with Reinforcement Learning on a Superconducting Processor18 Jun 2024 0 repositories listed
-
Reinforcement Learning for Corporate Bond Trading: A Sell Side Perspective18 Jun 2024 0 repositories listed
-
Physics-informed Imitative Reinforcement Learning for Real-world Driving18 Jun 2024 0 repositories listed
-
17 Jun 2024 0 repositories listed Syntology 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
Constructing Ancestral Recombination Graphs through Reinforcement Learning17 Jun 2024 0 repositories listed
-
Measuring memorization in RLHF for code completion17 Jun 2024 0 repositories listed
-
Optimal Transport-Assisted Risk-Sensitive Q-Learning17 Jun 2024 0 repositories listed
-
The Benefits of Power Regularization in Cooperative Reinforcement Learning17 Jun 2024 0 repositories listed
-
DIPPER: Direct Preference Optimization to Accelerate Primitive-Enabled Hierarchical Reinforcement Learning16 Jun 2024 0 repositories listed
-
InstructRL4Pix: Training Diffusion for Image Editing by Reinforcement Learning14 Jun 2024 0 repositories listed
-
CIMRL: Combining IMitation and Reinforcement Learning for Safe Autonomous Driving13 Jun 2024 0 repositories listed
-
Current applications and potential future directions of reinforcement learning-based Digital Twins in agriculture13 Jun 2024 0 repositories listed
-
DiffPoGAN: Diffusion Policies with Generative Adversarial Networks for Offline Reinforcement Learning13 Jun 2024 0 repositories listed
-
e-COP : Episodic Constrained Optimization of Policies13 Jun 2024 0 repositories listed
-
XLand-100B: A Large-Scale Multi-Task Dataset for In-Context Reinforcement Learning13 Jun 2024 0 repositories listed
-
Deep reinforcement learning with positional context for intraday trading12 Jun 2024 0 repositories listed
-
Explore-Go: Leveraging Exploration for Generalisation in Deep Reinforcement Learning12 Jun 2024 0 repositories listed
-
How social reinforcement learning can lead to metastable polarisation and the voter model12 Jun 2024 0 repositories listed
-
Optimizing Deep Reinforcement Learning for Adaptive Robotic Arm Control12 Jun 2024 0 repositories listed
-
Reinforcement Learning for High-Level Strategic Control in Tower Defense Games12 Jun 2024 0 repositories listed
-
RILe: Reinforced Imitation Learning12 Jun 2024 0 repositories listed
-
Time-Constrained Robust MDPs12 Jun 2024 0 repositories listed
-
CDSA: Conservative Denoising Score-based Algorithm for Offline Reinforcement Learning11 Jun 2024 0 repositories listed
-
Hybrid Reinforcement Learning from Offline Observation Alone11 Jun 2024 0 repositories listed
-
Learning Reward and Policy Jointly from Demonstration and Preference Improves Alignment11 Jun 2024 0 repositories listed
-
Boosting Robustness in Preference-Based Reinforcement Learning with Dynamic Sparsity10 Jun 2024 0 repositories listed
-
Risk Sensitivity in Markov Games and Multi-Agent Reinforcement Learning: A Systematic Review10 Jun 2024 0 repositories listed
-
Verification-Guided Shielding for Deep Reinforcement Learning10 Jun 2024 0 repositories listed
-
LGR2: Language Guided Reward Relabeling for Accelerating Hierarchical Reinforcement Learning9 Jun 2024 0 repositories listed
-
Enhanced Flight Envelope Protection: A Novel Reinforcement Learning Approach8 Jun 2024 0 repositories listed
-
Reinforcement Learning for Intensity Control: An Application to Choice-Based Network Revenue Management8 Jun 2024 0 repositories listed
-
AC4MPC: Actor-Critic Reinforcement Learning for Nonlinear Model Predictive Control6 Jun 2024 0 repositories listed
-
ATraDiff: Accelerating Online Reinforcement Learning with Imaginary Trajectories6 Jun 2024 0 repositories listed
-
Behavior-Targeted Attack on Reinforcement Learning with Limited Access to Victim's Policy6 Jun 2024 0 repositories listed
-
Bootstrapping Expectiles in Reinforcement Learning6 Jun 2024 0 repositories listed
-
Breeding Programs Optimization with Reinforcement Learning6 Jun 2024 0 repositories listed
-
Excluding the Irrelevant: Focusing Reinforcement Learning through Continuous Action Masking6 Jun 2024 0 repositories listed
-
Exploring Pessimism and Optimism Dynamics in Deep Reinforcement Learning6 Jun 2024 0 repositories listed
-
GenSafe: A Generalizable Safety Enhancer for Safe Reinforcement Learning Algorithms Based on Reduced Order Markov Decision Process Model6 Jun 2024 0 repositories listed
-
Optimizing Autonomous Driving for Safety: A Human-Centric Approach with LLM-Enhanced RLHF6 Jun 2024 0 repositories listed
-
Self-Play with Adversarial Critic: Provable and Scalable Offline Alignment for Language Models6 Jun 2024 0 repositories listed
-
Inductive Generalization in Reinforcement Learning from Specifications5 Jun 2024 0 repositories listed
-
Physics-Guided Actor-Critic Reinforcement Learning for Swimming in Turbulence5 Jun 2024 0 repositories listed
-
Representation Learning For Efficient Deep Multi-Agent Reinforcement Learning5 Jun 2024 0 repositories listed
-
A Unifying Framework for Action-Conditional Self-Predictive Reinforcement Learning4 Jun 2024 0 repositories listed
-
Adaptive Preference Scaling for Reinforcement Learning with Human Feedback4 Jun 2024 0 repositories listed
-
Algorithmic Collusion in Dynamic Pricing with Deep Reinforcement Learning4 Jun 2024 0 repositories listed
-
FightLadder: A Benchmark for Competitive Multi-Agent Reinforcement Learning4 Jun 2024 0 repositories listed
-
iQRL -- Implicitly Quantized Representations for Sample-efficient Reinforcement Learning4 Jun 2024 0 repositories listed
-
Rectifying Reinforcement Learning for Reward Matching4 Jun 2024 0 repositories listed
-
Reinforcement learning-based architecture search for quantum machine learning4 Jun 2024 0 repositories listed
-
Reinforcement Learning with Lookahead Information4 Jun 2024 0 repositories listed
-
Test-Time Regret Minimization in Meta Reinforcement Learning4 Jun 2024 0 repositories listed
-
A New View on Planning in Online Reinforcement Learning3 Jun 2024 0 repositories listed
-
Causal prompting model-based offline reinforcement learning3 Jun 2024 0 repositories listed
-
Deep Reinforcement Learning Behavioral Mode Switching Using Optimal Control Based on a Latent Space Objective3 Jun 2024 0 repositories listed
-
Deep reinforcement learning for weakly coupled MDP's with continuous actions3 Jun 2024 0 repositories listed
-
Multi-agent assignment via state augmented reinforcement learning3 Jun 2024 0 repositories listed
-
Multi-Agent Reinforcement Learning Meets Leaf Sequencing in Radiotherapy3 Jun 2024 0 repositories listed
-
Scalable Ensembling For Mitigating Reward Overoptimisation3 Jun 2024 0 repositories listed
-
A Digital Twin Framework for Reinforcement Learning with Real-Time Self-Improvement via Human Assistive Teleoperation2 Jun 2024 0 repositories listed
-
Learning to Play 7 Wonders Duel Without Human Supervision2 Jun 2024 0 repositories listed
-
Maximum-Entropy Regularized Decision Transformer with Reward Relabelling for Dynamic Recommendation2 Jun 2024 0 repositories listed
-
Multi-Dimensional Optimization for Text Summarization via Reinforcement Learning1 Jun 2024 0 repositories listed
-
Decision Mamba: Reinforcement Learning via Hybrid Selective Sequence Modeling31 May 2024 0 repositories listed
-
Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF31 May 2024 0 repositories listed
Syntology lines on 2 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.