Browse State-of-the-Art › reinforcement-learning › Papers, page 46
reinforcement-learning
Papers archive 2025-07-28
archive papers tagged: 13,427 · with a code link: 4,119 · where Syntology ran a sample: 1,165 (973 with a run with no instrument failure, 192 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (1,165 of 13,427 tagged: 973 with a run with no instrument failure, 192 where every run was a failure of Syntology's instrument)
Page 46 of 135: papers 4,501 to 4,600 of 13,427, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
MPO: An Efficient Post-Processing Framework for Mixing Diverse Preference Alignment25 Feb 2025 0 repositories listed
-
Survey on Strategic Mining in Blockchain: A Reinforcement Learning Approach24 Feb 2025 0 repositories listed
-
Towards Reinforcement Learning for Exploration of Speculative Execution Vulnerabilities24 Feb 2025 0 repositories listed
-
Exploring Sentiment Manipulation by LLM-Enabled Intelligent Trading Agents22 Feb 2025 0 repositories listed
-
Towards User-level Private Reinforcement Learning with Human Feedback22 Feb 2025 0 repositories listed
-
The Evolving Landscape of LLM- and VLM-Integrated Reinforcement Learning21 Feb 2025 0 repositories listed
-
Towards a Reward-Free Reinforcement Learning Framework for Vehicle Control21 Feb 2025 0 repositories listed
-
Causal Mean Field Multi-Agent Reinforcement Learning20 Feb 2025 0 repositories listed
-
Is Q-learning an Ill-posed Problem?20 Feb 2025 0 repositories listed
-
μRL: Discovering Transient Execution Vulnerabilities Using Reinforcement Learning20 Feb 2025 0 repositories listed
-
SPRIG: Stackelberg Perception-Reinforcement Learning with Internal Game Dynamics20 Feb 2025 0 repositories listed
-
Multi-Target Radar Search and Track Using Sequence-Capable Deep Reinforcement Learning19 Feb 2025 0 repositories listed
-
Uncertainty quantification for Markov chains with application to temporal difference learning19 Feb 2025 0 repositories listed
-
Continuous Learning Conversational AI: A Personalized Agent Framework via A2C Reinforcement Learning18 Feb 2025 0 repositories listed
-
Implicit Repair with Reinforcement Learning in Emergent Communication18 Feb 2025 0 repositories listed
-
Self-Supervised Transformers as Iterative Solution Improvers for Constraint Satisfaction18 Feb 2025 0 repositories listed
-
Theorem Prover as a Judge for Synthetic Data Generation18 Feb 2025 0 repositories listed
-
FitLight: Federated Imitation Learning for Plug-and-Play Autonomous Traffic Signal Control17 Feb 2025 0 repositories listed
-
Intelligent Mobile AI-Generated Content Services via Interactive Prompt Engineering and Dynamic Service Provisioning17 Feb 2025 0 repositories listed
-
Learning to Reason at the Frontier of Learnability17 Feb 2025 0 repositories listed
-
Theoretical Barriers in Bellman-Based Reinforcement Learning17 Feb 2025 0 repositories listed
-
Solving Online Resource-Constrained Scheduling for Follow-Up Observation in Astronomy: a Reinforcement Learning Approach16 Feb 2025 0 repositories listed
-
A Tutorial on LLM Reasoning: Relevant Methods behind ChatGPT o115 Feb 2025 0 repositories listed
-
Rule-Bottleneck Reinforcement Learning: Joint Explanation and Decision Optimization for Resource Allocation with Language Agents15 Feb 2025 0 repositories listed
-
Tackling the Zero-Shot Reinforcement Learning Loss Directly15 Feb 2025 0 repositories listed
-
Causal Information Prioritization for Efficient Reinforcement Learning14 Feb 2025 0 repositories listed
-
Combinatorial Reinforcement Learning with Preference Feedback14 Feb 2025 0 repositories listed
-
Do We Need to Verify Step by Step? Rethinking Process Supervision from a Theoretical Perspective14 Feb 2025 0 repositories listed
-
Dynamic Reinforcement Learning for Actors14 Feb 2025 0 repositories listed
-
Reinforcement Learning based Constrained Optimal Control: an Interpretable Reward Design14 Feb 2025 0 repositories listed
-
Reinforcement Learning in Strategy-Based and Atari Games: A Review of Google DeepMinds Innovations14 Feb 2025 0 repositories listed
-
Self-Consistent Model-based Adaptation for Visual Reinforcement Learning14 Feb 2025 0 repositories listed
-
A Survey of Reinforcement Learning for Optimization in Automation13 Feb 2025 0 repositories listed
-
Analysis of Off-Policy n-Step TD-Learning with Linear Function Approximation13 Feb 2025 0 repositories listed
-
Coupled Rendezvous and Docking Maneuver control of satellite using Reinforcement learning-based Adaptive Fixed-Time Sliding Mode Controller13 Feb 2025 0 repositories listed
-
Variable Stiffness for Robust Locomotion through Reinforcement Learning13 Feb 2025 0 repositories listed
-
Deep Reinforcement Learning-Based User Scheduling for Collaborative Perception12 Feb 2025 0 repositories listed
-
Provably Robust Federated Reinforcement Learning12 Feb 2025 0 repositories listed
-
A Survey of In-Context Reinforcement Learning11 Feb 2025 0 repositories listed
-
Exploratory Diffusion Model for Unsupervised Reinforcement Learning11 Feb 2025 0 repositories listed
-
Logarithmic Regret for Online KL-Regularized Reinforcement Learning11 Feb 2025 0 repositories listed
-
Optimal Actuator Attacks on Autonomous Vehicles Using Reinforcement Learning11 Feb 2025 0 repositories listed
-
PICTS: A Novel Deep Reinforcement Learning Approach for Dynamic P-I Control in Scanning Probe Microscopy11 Feb 2025 0 repositories listed
-
Polynomial-Time Approximability of Constrained Reinforcement Learning11 Feb 2025 0 repositories listed
-
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization11 Feb 2025 0 repositories listed
-
Advancing Autonomous VLM Agents via Variational Subgoal-Conditioned Reinforcement Learning11 Feb 2025 0 repositories listed
-
Intelligent Offloading in Vehicular Edge Computing: A Comprehensive Review of Deep Reinforcement Learning Approaches and Architectures10 Feb 2025 0 repositories listed
-
A Survey on Explainable Deep Reinforcement Learning8 Feb 2025 0 repositories listed
-
Real Time Control of Tandem-Wing Experimental Platform Using Concerto Reinforcement Learning8 Feb 2025 0 repositories listed
-
Sequential Stochastic Combinatorial Optimization Using Hierarchal Reinforcement Learning8 Feb 2025 0 repositories listed
-
Behavior-Regularized Diffusion Policy Optimization for Offline Reinforcement Learning7 Feb 2025 0 repositories listed
-
Agency Is Frame-Dependent6 Feb 2025 0 repositories listed
-
Behavioral Entropy-Guided Dataset Generation for Offline Reinforcement Learning6 Feb 2025 0 repositories listed
-
Fairness Aware Reinforcement Learning via Proximal Policy Optimization6 Feb 2025 0 repositories listed
-
Reinforcement Learning on Dyads to Enhance Medication Adherence6 Feb 2025 0 repositories listed
-
Double Distillation Network for Multi-Agent Reinforcement Learning5 Feb 2025 0 repositories listed
-
OmniRL: In-Context Reinforcement Learning by Large-Scale Meta-Training in Randomized Worlds5 Feb 2025 0 repositories listed
-
Teaching Language Models to Critique via Reinforcement Learning5 Feb 2025 0 repositories listed
-
CH-MARL: Constrained Hierarchical Multiagent Reinforcement Learning for Sustainable Maritime Logistics4 Feb 2025 0 repositories listed
-
DHP: Discrete Hierarchical Planning for Hierarchical Reinforcement Learning Agents4 Feb 2025 0 repositories listed
-
DIME:Diffusion-Based Maximum Entropy Reinforcement Learning4 Feb 2025 0 repositories listed
-
Policy-Guided Causal State Representation for Offline Reinforcement Learning Recommendation4 Feb 2025 0 repositories listed
-
ACECODER: Acing Coder RL via Automated Test-Case Synthesis3 Feb 2025 0 repositories listed
-
Competitive Programming with Large Reasoning Models3 Feb 2025 0 repositories listed
-
Exploratory Utility Maximization Problem with Tsallis Entropy3 Feb 2025 0 repositories listed
-
Generalized Lanczos method for systematic optimization of neural-network quantum states3 Feb 2025 0 repositories listed
-
Preference VLM: Leveraging VLMs for Scalable Preference-Based Reinforcement Learning3 Feb 2025 0 repositories listed
-
Process-Supervised Reinforcement Learning for Code Generation3 Feb 2025 0 repositories listed
-
Reinforcement Learning for Long-Horizon Interactive LLM Agents3 Feb 2025 0 repositories listed
-
Reinforcement Learning with Segment Feedback3 Feb 2025 0 repositories listed
-
The Differences Between Direct Alignment Algorithms are a Blur3 Feb 2025 0 repositories listed
-
Toward Task Generalization via Memory Augmentation in Meta-Reinforcement Learning3 Feb 2025 0 repositories listed
-
VR-Robo: A Real-to-Sim-to-Real Framework for Visual Robot Navigation and Locomotion3 Feb 2025 0 repositories listed
-
Compositional Concept-Based Neuron-Level Interpretability for Deep Reinforcement Learning2 Feb 2025 0 repositories listed
-
Decorrelated Soft Actor-Critic for Efficient Deep Reinforcement Learning31 Jan 2025 0 repositories listed
-
Shaping Sparse Rewards in Reinforcement Learning: A Semi-supervised Approach31 Jan 2025 0 repositories listed
-
SpikingSoft: A Spiking Neuron Controller for Bio-inspired Locomotion with Soft Snake Robots31 Jan 2025 0 repositories listed
-
Hybrid Group Relative Policy Optimization: A Multi-Sample Approach to Enhancing Policy Optimization30 Jan 2025 0 repositories listed
-
Certificated Actor-Critic: Hierarchical Reinforcement Learning with Control Barrier Functions for Safe Navigation29 Jan 2025 0 repositories listed
-
Digital Twin Synchronization: Bridging the Sim-RL Agent to a Real-Time Robotic Additive Manufacturing Control29 Jan 2025 0 repositories listed
-
Reinforcement-Learning Portfolio Allocation with Dynamic Embedding of Market Information29 Jan 2025 0 repositories listed
-
The M-factor: A Novel Metric for Evaluating Neural Architecture Search in Resource-Constrained Environments29 Jan 2025 0 repositories listed
-
Applying Ensemble Models based on Graph Neural Network and Reinforcement Learning for Wind Power Forecasting28 Jan 2025 0 repositories listed
-
Improving Vision-Language-Action Model with Online Reinforcement Learning28 Jan 2025 0 repositories listed
-
Induced Modularity and Community Detection for Functionally Interpretable Reinforcement Learning28 Jan 2025 0 repositories listed
-
On the Interplay Between Sparsity and Training in Deep Reinforcement Learning28 Jan 2025 0 repositories listed
-
Safe Reinforcement Learning for Real-World Engine Control28 Jan 2025 0 repositories listed
-
Generative AI for Lyapunov Optimization Theory in UAV-based Low-Altitude Economy Networking27 Jan 2025 0 repositories listed
-
Reinforcement Learning for Quantum Circuit Design: Using Matrix Representations27 Jan 2025 0 repositories listed
-
Selective Experience Sharing in Reinforcement Learning Enhances Interference Management27 Jan 2025 0 repositories listed
-
27 Jan 2025 0 repositories listed Syntology 5 ran (of which 4 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified (of 9 harvested samples) · 4 pointer-only (licence)
-
Advancing TDFN: Precise Fixation Point Generation Using Reconstruction Differences26 Jan 2025 0 repositories listed
-
Contextual Knowledge Sharing in Multi-Agent Reinforcement Learning with Decentralized Communication and Coordination26 Jan 2025 0 repositories listed
-
Data Center Cooling System Optimization Using Offline Reinforcement Learning25 Jan 2025 0 repositories listed
-
Extensive Exploration in Complex Traffic Scenarios using Hierarchical Reinforcement Learning25 Jan 2025 0 repositories listed
-
Faster Machine Translation Ensembling with Reinforcement Learning and Competitive Correction25 Jan 2025 0 repositories listed
-
Music Generation using Human-In-The-Loop Reinforcement Learning25 Jan 2025 0 repositories listed
-
Predictive Lagrangian Optimization for Constrained Reinforcement Learning25 Jan 2025 0 repositories listed
-
Reinforcement Learning for Efficient Returns Management24 Jan 2025 0 repositories listed
Syntology lines on 2 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.