Methods › Reinforcement Learning › Policy Gradient Methods › PPO › Papers, page 2
Proximal Policy Optimization
PPO
Papers archive 2025-07-28
archive papers tagged: 949 · with a code link: 397 · where Syntology ran a sample: 139 (114 with a run with no instrument failure, 25 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (139 of 949 tagged: 114 with a run with no instrument failure, 25 where every run was a failure of Syntology's instrument)
Page 2 of 10: papers 101 to 200 of 949, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
A Comprehensive LLM-powered Framework for Driving Intelligence Evaluation 7 Mar 2025 · 2 repositories · arXiv:2503.05164
-
BEVDriver: Leveraging BEV Maps in LLMs for Robust Closed-Loop Driving 5 Mar 2025 · 1 repository · arXiv:2503.03074
-
Perceptual Motor Learning with Active Inference Framework for Robust Lateral Control 3 Mar 2025 · 0 repositories · arXiv:2503.01676
-
CAPS: Context-Aware Priority Sampling for Enhanced Imitation Learning in Autonomous Driving 3 Mar 2025 · 0 repositories · arXiv:2503.01650
-
What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret 3 Mar 2025 · 0 repositories · arXiv:2503.01491
-
CARIL: Confidence-Aware Regression in Imitation Learning for Autonomous Driving 2 Mar 2025 · 1 repository · arXiv:2503.00783
-
Multimodal Dreaming: A Global Workspace Approach to World Model-Based Reinforcement Learning 28 Feb 2025 · 0 repositories · arXiv:2502.21142
-
Shared Autonomy for Proximal Teaching 27 Feb 2025 · 0 repositories · arXiv:2502.19899
-
SoRFT: Issue Resolving with Subtask-oriented Reinforced Fine-Tuning 27 Feb 2025 · 0 repositories · arXiv:2502.20127
-
Lean and Mean: Decoupled Value Policy Optimization with Global Value Guidance 24 Feb 2025 · 0 repositories · arXiv:2502.16944
-
Ensemble RL through Classifier Models: Enhancing Risk-Return Trade-offs in Trading Strategies 23 Feb 2025 · 0 repositories · arXiv:2502.17518
-
SPRIG: Stackelberg Perception-Reinforcement Learning with Internal Game Dynamics 20 Feb 2025 · 0 repositories · arXiv:2502.14264
-
Synth It Like KITTI: Synthetic Data Generation for Object Detection in Driving Scenarios 20 Feb 2025 · 1 repository · arXiv:2502.15076
-
Sce2DriveX: A Generalized MLLM Framework for Scene-to-Drive Learning 19 Feb 2025 · 0 repositories · arXiv:2502.14917
-
Hovering Flight of Soft-Actuated Insect-Scale Micro Aerial Vehicles using Deep Reinforcement Learning 17 Feb 2025 · 0 repositories · arXiv:2502.12355
-
MC-BEVRO: Multi-Camera Bird Eye View Road Occupancy Detection for Traffic Monitoring 16 Feb 2025 · 0 repositories · arXiv:2502.11287
-
Reevaluating Policy Gradient Methods for Imperfect-Information Games 13 Feb 2025 · 2 repositories · arXiv:2502.08938Syntology official (archive's flag): 1 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Salience-Invariant Consistent Policy Learning for Generalization in Visual Reinforcement Learning 12 Feb 2025 · 0 repositories · arXiv:2502.08336
-
On the Emergence of Thinking in LLMs I: Searching for the Right Intuition 10 Feb 2025 · 4 repositories · arXiv:2502.06773Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 13 harvested samples) · 9 pointer-only (licence)
-
Fairness Aware Reinforcement Learning via Proximal Policy Optimization 6 Feb 2025 · 0 repositories · arXiv:2502.03953
-
Label Anything: An Interpretable, High-Fidelity and Prompt-Free Annotator 5 Feb 2025 · 0 repositories · arXiv:2502.02972
-
Regret-Optimized Portfolio Enhancement through Deep Reinforcement Learning and Future Looking Rewards 4 Feb 2025 · 0 repositories · arXiv:2502.02619
-
Mitigation of Camouflaged Adversarial Attacks in Autonomous Vehicles--A Case Study Using CARLA Simulator 3 Feb 2025 · 0 repositories · arXiv:2502.05208
-
SynthmanticLiDAR: A Synthetic Dataset for Semantic Segmentation on LiDAR Imaging 31 Jan 2025 · 1 repository · arXiv:2501.19035
-
The Energy Loss Phenomenon in RLHF: A New Perspective on Mitigating Reward Hacking 31 Jan 2025 · 0 repositories · arXiv:2501.19358
-
Hybrid Group Relative Policy Optimization: A Multi-Sample Approach to Enhancing Policy Optimization 30 Jan 2025 · 0 repositories · arXiv:2502.01652
-
SSF-PAN: Semantic Scene Flow-Based Perception for Autonomous Navigation in Traffic Scenarios 28 Jan 2025 · 0 repositories · arXiv:2501.16754
-
EvoRL: A GPU-accelerated Framework for Evolutionary Reinforcement Learning 25 Jan 2025 · 2 repositories · arXiv:2501.15129
-
AdaWM: Adaptive World Model based Planning for Autonomous Driving 22 Jan 2025 · 0 repositories · arXiv:2501.13072
-
HEPPO: Hardware-Efficient Proximal Policy Optimization -- A Universal Pipelined Architecture for Generalized Advantage Estimation 22 Jan 2025 · 0 repositories · arXiv:2501.12703
-
To Measure or Not: A Cost-Sensitive, Selective Measuring Environment for Agricultural Management Decisions with Reinforcement Learning 22 Jan 2025 · 1 repository · arXiv:2501.12823
-
Explainable Lane Change Prediction for Near-Crash Scenarios Using Knowledge Graph Embeddings and Retrieval Augmented Generation 20 Jan 2025 · 0 repositories · arXiv:2501.11560
-
Classical and Deep Reinforcement Learning Inventory Control Policies for Pharmaceutical Supply Chains with Perishability and Non-Stationarity 18 Jan 2025 · 0 repositories · arXiv:2501.10895
-
LeapVAD: A Leap in Autonomous Driving via Cognitive Perception and Dual-Process Thinking 14 Jan 2025 · 1 repository · arXiv:2501.08168Syntology 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 1 pointer-only (licence)
-
Optimization of Link Configuration for Satellite Communication Using Reinforcement Learning 14 Jan 2025 · 0 repositories · arXiv:2501.08220
-
A Hybrid Framework for Reinsurance Optimization: Integrating Generative Models and Reinforcement Learning 11 Jan 2025 · 1 repository · arXiv:2501.06404
-
CuRLA: Curriculum Learning Based Deep Reinforcement Learning for Autonomous Driving 9 Jan 2025 · 0 repositories · arXiv:2501.04982
-
LearningFlow: Automated Policy Learning Workflow for Urban Driving with Large Language Models 9 Jan 2025 · 0 repositories · arXiv:2501.05057
-
REINFORCE++: A Simple and Efficient Approach for Aligning Large Language Models 4 Jan 2025 · 5 repositories · arXiv:2501.03262
-
DriveGPT4-V2: Harnessing Large Language Model Capabilities for Enhanced Closed-Loop Autonomous Driving 1 Jan 2025 · 0 repositories
-
Plug-and-Play PPO: An Adaptive Point Prompt Optimizer Making SAM Greater 1 Jan 2025 · 1 repository
-
VLMs-Guided Representation Distillation for Efficient Vision-Based Reinforcement Learning 1 Jan 2025 · 0 repositories
-
RFPPO: Motion Dynamic RRT based Fluid Field - PPO for Dynamic TF/TA Routing Planning 28 Dec 2024 · 0 repositories · arXiv:2412.20098
-
Graph-attention-based Casual Discovery with Trust Region-navigated Clipping Policy Optimization 27 Dec 2024 · 0 repositories · arXiv:2412.19578
-
Preventive Energy Management for Distribution Systems Under Uncertain Events: A Deep Reinforcement Learning Approach 26 Dec 2024 · 0 repositories · arXiv:2412.19382
-
Deep Learning-Based Traffic-Aware Base Station Sleep Mode and Cell Zooming Strategy in RIS-Aided Multi-Cell Networks 25 Dec 2024 · 0 repositories · arXiv:2412.18983
-
Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization 24 Dec 2024 · 0 repositories · arXiv:2412.18279
-
VLM-RL: A Unified Vision Language Models and Reinforcement Learning Framework for Safe Autonomous Driving 20 Dec 2024 · 0 repositories · arXiv:2412.15544
-
Active Reinforcement Learning Strategies for Offline Policy Improvement 17 Dec 2024 · 0 repositories · arXiv:2412.13106
-
Efficient Policy Adaptation with Contrastive Prompt Ensemble for Embodied Agents 16 Dec 2024 · 0 repositories · arXiv:2412.11484Syntology 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample) · 1 pointer-only (licence)
-
Automated Driving with Evolution Capability: A Reinforcement Learning Method with Monotonic Performance Enhancement 14 Dec 2024 · 0 repositories · arXiv:2412.10822
-
EI-Drive: A Platform for Cooperative Perception with Realistic Communication Models 13 Dec 2024 · 0 repositories · arXiv:2412.09782
-
TrafficLoc: Localizing Traffic Surveillance Cameras in 3D Scenes 13 Dec 2024 · 0 repositories · arXiv:2412.10308
-
WiseAD: Knowledge Augmented End-to-End Autonomous Driving with Vision-Language Model 13 Dec 2024 · 1 repository · arXiv:2412.09951Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 1 honoured, 0 violated, 2 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
eCARLA-scenes: A synthetically generated dataset for event-based optical flow prediction 12 Dec 2024 · 2 repositories · arXiv:2412.09209
-
Hidden Biases of End-to-End Driving Datasets 12 Dec 2024 · 1 repository · arXiv:2412.09602Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 2 honoured, 1 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Bench2Drive-R: Turning Real World Data into Reactive Closed-Loop Autonomous Driving Benchmark by Generative Model 11 Dec 2024 · 0 repositories · arXiv:2412.09647
-
A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning 10 Dec 2024 · 1 repository · arXiv:2412.07165Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
A Scalable Decentralized Reinforcement Learning Framework for UAV Target Localization Using Recurrent PPO 9 Dec 2024 · 0 repositories · arXiv:2412.06231
-
Traffic Co-Simulation Framework Empowered by Infrastructure Camera Sensing and Reinforcement Learning 5 Dec 2024 · 0 repositories · arXiv:2412.03925
-
Using Cooperative Co-evolutionary Search to Generate Metamorphic Test Cases for Autonomous Driving Systems 5 Dec 2024 · 0 repositories · arXiv:2412.03843
-
FedPAW: Federated Learning with Personalized Aggregation Weights for Urban Vehicle Speed Prediction 2 Dec 2024 · 1 repository · arXiv:2412.01281
-
HPRM: High-Performance Robotic Middleware for Intelligent Autonomous Systems 2 Dec 2024 · 0 repositories · arXiv:2412.01799
-
Realistic Corner Case Generation for Autonomous Vehicles with Multimodal Large Language Model 29 Nov 2024 · 0 repositories · arXiv:2412.00243
-
Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards 26 Nov 2024 · 0 repositories · arXiv:2411.17861
-
From Dashcam Videos to Driving Simulations: Stress Testing Automated Vehicles against Rare Events 25 Nov 2024 · 0 repositories · arXiv:2411.16027
-
Generating Out-Of-Distribution Scenarios Using Language Models 25 Nov 2024 · 0 repositories · arXiv:2411.16554
-
Imperceptible Adversarial Examples in the Physical World 25 Nov 2024 · 0 repositories · arXiv:2411.16622
-
SynDiff-AD: Improving Semantic Segmentation and End-to-End Autonomous Driving with Synthetic Data from Latent Diffusion Models 25 Nov 2024 · 0 repositories · arXiv:2411.16776
-
Reward Fine-Tuning Two-Step Diffusion Models via Learning Differentiable Latent-Space Surrogate Reward 22 Nov 2024 · 0 repositories · arXiv:2411.15247
-
Provably Efficient Action-Manipulation Attack Against Continuous Reinforcement Learning 20 Nov 2024 · 0 repositories · arXiv:2411.13116
-
WHALES: A Multi-agent Scheduling Dataset for Enhanced Cooperation in Autonomous Driving 20 Nov 2024 · 1 repository · arXiv:2411.13340
-
Fast Convergence of Softmax Policy Mirror Ascent 18 Nov 2024 · 0 repositories · arXiv:2411.12042
-
Dynamics of Resource Allocation in O-RANs: An In-depth Exploration of On-Policy and Off-Policy Deep Reinforcement Learning for Real-Time Applications 17 Nov 2024 · 0 repositories · arXiv:2412.01839
-
Imagine-2-Drive: Leveraging High-Fidelity World Models via Multi-Modal Diffusion Policies 15 Nov 2024 · 0 repositories · arXiv:2411.10171
-
Edge Caching Optimization with PPO and Transfer Learning for Dynamic Environments 14 Nov 2024 · 0 repositories · arXiv:2411.09812
-
Innate-Values-driven Reinforcement Learning based Cognitive Modeling 14 Nov 2024 · 0 repositories · arXiv:2411.09160
-
Exploring Multi-Agent Reinforcement Learning for Unrelated Parallel Machine Scheduling 12 Nov 2024 · 0 repositories · arXiv:2411.07634
-
Research on reinforcement learning based warehouse robot navigation algorithm in complex warehouse layout 9 Nov 2024 · 0 repositories · arXiv:2411.06128
-
Querying Perception Streams with Spatial Regular Expressions 8 Nov 2024 · 0 repositories · arXiv:2411.05946
-
Think Smart, Act SMARL! Analyzing Probabilistic Logic Shields for Multi-Agent Reinforcement Learning 7 Nov 2024 · 1 repository · arXiv:2411.04867
-
A Comparative Study of Deep Reinforcement Learning for Crop Production Management 6 Nov 2024 · 0 repositories · arXiv:2411.04106
-
Federated Data-Driven Kalman Filtering for State Estimation 6 Nov 2024 · 0 repositories · arXiv:2411.05847
-
SALSA: Soup-based Alignment Learning for Stronger Adaptation in RLHF 4 Nov 2024 · 0 repositories · arXiv:2411.01798
-
Beyond the Boundaries of Proximal Policy Optimization 1 Nov 2024 · 0 repositories · arXiv:2411.00666
-
CALE: Continuous Arcade Learning Environment 31 Oct 2024 · 1 repository · arXiv:2410.23810Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Optical Lens Attack on Monocular Depth Estimation for Autonomous Driving 31 Oct 2024 · 0 repositories · arXiv:2411.00192
-
MDCure: A Scalable Pipeline for Multi-Document Instruction-Following 30 Oct 2024 · 1 repository · arXiv:2410.23463
-
Self-Driving Car Racing: Application of Deep Reinforcement Learning 30 Oct 2024 · 0 repositories · arXiv:2410.22766
-
Pre-Trained Vision Models as Perception Backbones for Safety Filters in Autonomous Driving 29 Oct 2024 · 0 repositories · arXiv:2410.22585
-
SS3DM: Benchmarking Street-View Surface Reconstruction with a Synthetic 3D Mesh Dataset 29 Oct 2024 · 0 repositories · arXiv:2410.21739
-
Capacity-Aware Planning and Scheduling in Budget-Constrained Monotonic MDPs: A Meta-RL Approach 28 Oct 2024 · 0 repositories · arXiv:2410.21249
-
Getting By Goal Misgeneralization With a Little Help From a Mentor 28 Oct 2024 · 0 repositories · arXiv:2410.21052
-
Fast Best-of-N Decoding via Speculative Rejection 26 Oct 2024 · 1 repository · arXiv:2410.20290Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Who is Responsible? Explaining Safety Violations in Multi-Agent Cyber-Physical Systems 26 Oct 2024 · 0 repositories · arXiv:2410.20288
-
Aligning CodeLLMs with Direct Preference Optimization 24 Oct 2024 · 0 repositories · arXiv:2410.18585
-
CARLA2Real: a tool for reducing the sim2real gap in CARLA simulator 23 Oct 2024 · 1 repository · arXiv:2410.18238
-
Exploring RL-based LLM Training for Formal Language Tasks with Programmed Rewards 22 Oct 2024 · 1 repository · arXiv:2410.17126
-
Hierarchical Multi-agent Reinforcement Learning for Cyber Network Defense 22 Oct 2024 · 0 repositories · arXiv:2410.17351
-
How to Build a Pre-trained Multimodal model for Simultaneously Chatting and Decision-making? 21 Oct 2024 · 0 repositories · arXiv:2410.15885