Browse State-of-the-Art › Offline RL › Papers, page 4
Offline RL
Papers archive 2025-07-28
archive papers tagged: 755 · with a code link: 310 · where Syntology ran a sample: 164 (139 with a run with no instrument failure, 25 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (164 of 755 tagged: 139 with a run with no instrument failure, 25 where every run was a failure of Syntology's instrument)
Page 4 of 8: papers 301 to 400 of 755, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
26 Dec 2020 1 repository listed
-
21 Dec 2020 1 repository listed
-
1 Dec 2020 1 repository listed
-
22 Oct 2020 1 repository listed
-
12 Oct 2020 1 repository listed Syntology official: no sample here; runs from other or unrecorded repositories · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
2 Oct 2020 1 repository listed Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 4 harvested samples) · 3 pointer-only (licence)
-
1 Jan 2020 1 repository listed
-
26 Nov 2019 1 repository listed
-
10 Jul 2019 1 repository listed
-
From Novelty to Imitation: Self-Distilled Rewards for Offline Reinforcement Learning17 Jul 2025 0 repositories listed
-
Robust Bandwidth Estimation for Real-Time Communication with Offline Reinforcement Learning8 Jul 2025 0 repositories listed
-
26 Jun 2025 0 repositories listed Syntology 3 ran (of which 2 constructed an object rather than computing a result; 3 with no instrument failure: 1 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Optimal Single-Policy Sample Complexity and Transient Coverage for Average-Reward Offline RL26 Jun 2025 0 repositories listed
-
IntelliLung: Advancing Safe Mechanical Ventilation using Offline RL with Hybrid Actions and Clinically Aligned Rewards17 Jun 2025 0 repositories listed
-
Toward Explainable Offline RL: Analyzing Representations in Intrinsically Motivated Decision Transformers16 Jun 2025 0 repositories listed
-
MOORL: A Framework for Integrating Offline-Online Reinforcement Learning11 Jun 2025 0 repositories listed
-
How to Provably Improve Return Conditioned Supervised Learning?10 Jun 2025 0 repositories listed
-
Policy-Based Trajectory Clustering in Offline Reinforcement Learning10 Jun 2025 0 repositories listed
-
Semi-gradient DICE for Offline Constrained Reinforcement Learning10 Jun 2025 0 repositories listed
-
Accelerating Diffusion Models in Offline RL via Reward-Aware Consistency Trajectory Distillation9 Jun 2025 0 repositories listed
-
Learning to Clarify by Reinforcement Learning Through Reward-Weighted Fine-Tuning8 Jun 2025 0 repositories listed
-
ADG: Ambient Diffusion-Guided Dataset Recovery for Corruption-Robust Offline Reinforcement Learning29 May 2025 0 repositories listed
-
Enhanced DACER Algorithm with High Diffusion Efficiency29 May 2025 0 repositories listed
-
Scaling Offline RL via Efficient and Expressive Shortcut Models28 May 2025 0 repositories listed
-
GenPO: Generative Diffusion Models Meet On-Policy Reinforcement Learning24 May 2025 0 repositories listed
-
Diffusion Self-Weighted Guidance for Offline Reinforcement Learning23 May 2025 0 repositories listed
-
Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only22 May 2025 0 repositories listed
-
Offline Guarded Safe Reinforcement Learning for Medical Treatment Optimization Strategies22 May 2025 0 repositories listed
-
Unearthing Gems from Stones: Policy Optimization with Negative Sample Augmentation for LLM Reasoning20 May 2025 0 repositories listed
-
Your Offline Policy is Not Trustworthy: Bilevel Reinforcement Learning for Sequential Portfolio Optimization19 May 2025 0 repositories listed
-
Prior-Guided Diffusion Planning for Offline Reinforcement Learning16 May 2025 0 repositories listed
-
Reinforcement Learning for Individual Optimal Policy from Heterogeneous Data14 May 2025 0 repositories listed
-
Feasibility-Aware Pessimistic Estimation: Toward Long-Horizon Safety in Offline RL13 May 2025 0 repositories listed
-
Cache-Efficient Posterior Sampling for Reinforcement Learning with LLM-Derived Priors Across Discrete and Continuous Domains12 May 2025 0 repositories listed
-
What Matters for Batch Online Reinforcement Learning in Robotics?12 May 2025 0 repositories listed
-
Video-Enhanced Offline Reinforcement Learning: A Model-Based Approach10 May 2025 0 repositories listed
-
Pretraining a Shared Q-Network for Data-Efficient Offline Reinforcement Learning9 May 2025 0 repositories listed
-
Taming OOD Actions for Offline Reinforcement Learning: An Advantage-Based Approach8 May 2025 0 repositories listed
-
Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study4 May 2025 0 repositories listed
-
Analytic Energy-Guided Policy Optimization for Offline Reinforcement Learning3 May 2025 0 repositories listed
-
Offline Robotic World Model: Learning Robotic Policies without a Physics Simulator23 Apr 2025 0 repositories listed
-
16 Apr 2025 0 repositories listed Syntology 10 ran (of which 9 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified (of 14 harvested samples) · 14 pointer-only (licence)
-
Towards Optimal Differentially Private Regret Bounds in Linear MDPs12 Apr 2025 0 repositories listed
-
Decision SpikeFormer: Spike-Driven Transformer for Decision Making4 Apr 2025 0 repositories listed
-
Model-Based Offline Reinforcement Learning with Adversarial Data Augmentation26 Mar 2025 0 repositories listed
-
Offline Reinforcement Learning with Discrete Diffusion Skills26 Mar 2025 0 repositories listed
-
Behaviour Discovery and Attribution for Explainable Reinforcement Learning19 Mar 2025 0 repositories listed
-
Evaluation-Time Policy Switching for Offline Reinforcement Learning15 Mar 2025 0 repositories listed
-
The Pitfalls of Imitation Learning when Actions are Continuous12 Mar 2025 0 repositories listed
-
Policy Regularization on Globally Accessible States in Cross-Dynamics Reinforcement Learning10 Mar 2025 0 repositories listed
-
Energy-Weighted Flow Matching for Offline Reinforcement Learning6 Mar 2025 0 repositories listed
-
Yes, Q-learning Helps Offline In-Context RL24 Feb 2025 0 repositories listed
-
Enhancing Offline Model-Based RL via Active Model Selection: A Bayesian Optimization Perspective17 Feb 2025 0 repositories listed
-
Which Features are Best for Successor Features?15 Feb 2025 0 repositories listed
-
Diverse Transformer Decoding for Offline Reinforcement Learning Using Financial Algorithmic Approaches13 Feb 2025 0 repositories listed
-
Enhancing Pre-Trained Decision Transformers with Prompt-Tuning Bandits7 Feb 2025 0 repositories listed
-
Behavioral Entropy-Guided Dataset Generation for Offline Reinforcement Learning6 Feb 2025 0 repositories listed
-
OmniRL: In-Context Reinforcement Learning by Large-Scale Meta-Training in Randomized Worlds5 Feb 2025 0 repositories listed
-
Policy-Guided Causal State Representation for Offline Reinforcement Learning Recommendation4 Feb 2025 0 repositories listed
-
Resilient UAV Trajectory Planning via Few-Shot Meta-Offline Reinforcement Learning3 Feb 2025 0 repositories listed
-
Flexible Blood Glucose Control: Offline Reinforcement Learning from Human Feedback27 Jan 2025 0 repositories listed
-
Data Center Cooling System Optimization Using Offline Reinforcement Learning25 Jan 2025 0 repositories listed
-
Large Language Model driven Policy Exploration for Recommender Systems23 Jan 2025 0 repositories listed
-
DRDT3: Diffusion-Refined Decision Test-Time Training Model12 Jan 2025 0 repositories listed
-
SR-Reward: Taking The Path More Traveled4 Jan 2025 0 repositories listed
-
On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures3 Jan 2025 0 repositories listed
-
Goal-Conditioned Data Augmentation for Offline Reinforcement Learning29 Dec 2024 0 repositories listed
-
Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization24 Dec 2024 0 repositories listed
-
AdaCred: Adaptive Causal Decision Transformers with Feature Crediting19 Dec 2024 0 repositories listed
-
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone9 Dec 2024 0 repositories listed
-
Finer Behavioral Foundation Models via Auto-Regressive Features and Advantage Weighting5 Dec 2024 0 repositories listed
-
Improving Dynamic Object Interactions in Text-to-Video Generation with AI Feedback3 Dec 2024 0 repositories listed
-
27 Nov 2024 0 repositories listed Syntology 2 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; every one of the 2 samples that ran constructed an object rather than computing a result (of 2 harvested samples)
-
LLM-Based Offline Learning for Embodied Agents via Consistency-Guided Reward Ensemble26 Nov 2024 0 repositories listed
-
PROGRESSOR: A Perceptually Guided Reward Estimator with Self-Supervised Online Refinement26 Nov 2024 0 repositories listed
-
Preserving Expert-Level Privacy in Offline Reinforcement Learning18 Nov 2024 0 repositories listed
-
Navigation with QPHIL: Quantizing Planner for Hierarchical Implicit Q-Learning12 Nov 2024 0 repositories listed
-
Streetwise Agents: Empowering Offline RL Policies to Outsmart Exogenous Stochastic Disturbances in RTC11 Nov 2024 0 repositories listed
-
OffLight: An Offline Multi-Agent Reinforcement Learning Framework for Traffic Signal Control10 Nov 2024 0 repositories listed
-
Real-World Offline Reinforcement Learning from Vision Language Model Feedback8 Nov 2024 0 repositories listed
-
Q-SFT: Q-Learning for Language Models via Supervised Fine-Tuning7 Nov 2024 0 repositories listed
-
Offline Reinforcement Learning and Sequence Modeling for Downlink Link Adaptation30 Oct 2024 0 repositories listed
-
Offline reinforcement learning for job-shop scheduling problems21 Oct 2024 0 repositories listed
-
Solving Continual Offline RL through Selective Weights Activation on Aligned Spaces21 Oct 2024 0 repositories listed
-
Off-dynamics Conditional Diffusion Planners16 Oct 2024 0 repositories listed
-
DIAR: Diffusion-model-guided Implicit Q-learning with Adaptive Revaluation15 Oct 2024 0 repositories listed
-
Diffusion-Based Offline RL for Improved Decision-Making in Augmented ARC Task15 Oct 2024 0 repositories listed
-
Multi-Objective-Optimization Multi-AUV Assisted Data Collection Framework for IoUT Based on Offline Reinforcement Learning15 Oct 2024 0 repositories listed
-
Integrating Reinforcement Learning and Large Language Models for Crop Production Process Management Optimization and Control through A New Knowledge-Based Deep Learning Paradigm13 Oct 2024 0 repositories listed
-
Offline Inverse Constrained Reinforcement Learning for Safe-Critical Decision Making in Healthcare10 Oct 2024 0 repositories listed
-
ComaDICE: Offline Cooperative Multi-Agent Reinforcement Learning with Stationary Distribution Shift Regularization2 Oct 2024 0 repositories listed
-
The Smart Buildings Control Suite: A Diverse Open Source Benchmark to Evaluate and Scale HVAC Control Policies for Sustainability2 Oct 2024 0 repositories listed
-
OffRIPP: Offline RL-based Informative Path Planning25 Sep 2024 0 repositories listed
-
Development and Validation of Heparin Dosing Policies Using an Offline Reinforcement Learning Algorithm24 Sep 2024 0 repositories listed
-
KAN v.s. MLP for Offline Reinforcement Learning15 Sep 2024 0 repositories listed
-
Q-value Regularized Decision ConvFormer for Offline Reinforcement Learning12 Sep 2024 0 repositories listed
-
Enhancing Cross-domain Pre-Trained Decision Transformers with Adaptive Attention11 Sep 2024 0 repositories listed
-
Tractable Offline Learning of Regular Decision Processes4 Sep 2024 0 repositories listed
-
Skills Regularized Task Decomposition for Multi-task Offline Reinforcement Learning28 Aug 2024 0 repositories listed
Syntology lines on 6 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.