Methods › Reinforcement Learning › Policy Gradient Methods › PPO › Papers, page 5
Proximal Policy Optimization
PPO
Papers archive 2025-07-28
archive papers tagged: 949 · with a code link: 397 · where Syntology ran a sample: 139 (114 with a run with no instrument failure, 25 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (139 of 949 tagged: 114 with a run with no instrument failure, 25 where every run was a failure of Syntology's instrument)
Page 5 of 10: papers 401 to 500 of 949, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Safety Aware Autonomous Path Planning Using Model Predictive Reinforcement Learning for Inland Waterways 16 Nov 2023 · 0 repositories · arXiv:2311.09878
-
Clipped-Objective Policy Gradients for Pessimistic Policy Optimization 10 Nov 2023 · 1 repository · arXiv:2311.05846
-
Black-Box Prompt Optimization: Aligning Large Language Models without Model Training 7 Nov 2023 · 1 repository · arXiv:2311.04155Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples)
-
Epidemic Decision-making System Based Federated Reinforcement Learning 3 Nov 2023 · 0 repositories · arXiv:2311.01749
-
Dropout Strategy in Reinforcement Learning: Limiting the Surrogate Objective Variance in Policy Optimization Methods 31 Oct 2023 · 0 repositories · arXiv:2310.20380
-
Bird's Eye View Based Pretrained World model for Visual Navigation 28 Oct 2023 · 0 repositories · arXiv:2310.18847
-
Reward Scale Robustness for Proximal Policy Optimization via DreamerV3 Tricks 26 Oct 2023 · 0 repositories · arXiv:2310.17805Syntology 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 1 pointer-only (licence)
-
CycleAlign: Iterative Distillation from Black-box LLM to White-box Models for Better Human Alignment 25 Oct 2023 · 0 repositories · arXiv:2310.16271
-
SuperHF: Supervised Iterative Learning from Human Feedback 25 Oct 2023 · 1 repository · arXiv:2310.16763Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
Safe Navigation: Training Autonomous Vehicles using Deep Reinforcement Learning in CARLA 23 Oct 2023 · 1 repository · arXiv:2311.10735Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
Automatic Unit Test Data Generation and Actor-Critic Reinforcement Learning for Code Synthesis 20 Oct 2023 · 2 repositories · arXiv:2310.13669
-
LeTFuser: Light-weight End-to-end Transformer-Based Sensor Fusion for Autonomous Driving with Multi-Task Learning 19 Oct 2023 · 1 repository · arXiv:2310.13135
-
ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models 16 Oct 2023 · 3 repositories · arXiv:2310.10505Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 13 harvested samples) · 13 pointer-only (licence)
-
Optimizing the Placement of Roadside LiDARs for Autonomous Driving 11 Oct 2023 · 0 repositories · arXiv:2310.07247
-
Distributional Soft Actor-Critic with Three Refinements 9 Oct 2023 · 2 repositories · arXiv:2310.05858Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Reinforcement Learning in the Era of LLMs: What is Essential? What is needed? An RL Perspective on RLHF, Prompting, and Beyond 9 Oct 2023 · 2 repositories · arXiv:2310.06147Syntology 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
DialCoT Meets PPO: Decomposing and Exploring Reasoning Paths in Smaller Language Models 8 Oct 2023 · 1 repository · arXiv:2310.05074
-
FP3O: Enabling Proximal Policy Optimization in Multi-Agent Cooperation with Parameter-Sharing Versatility 8 Oct 2023 · 0 repositories · arXiv:2310.05053
-
Proximal Policy Optimization-Based Reinforcement Learning Approach for DC-DC Boost Converter Control: A Comparative Evaluation Against Traditional Control Techniques 4 Oct 2023 · 0 repositories · arXiv:2310.02945
-
Reward Model Ensembles Help Mitigate Overoptimization 4 Oct 2023 · 2 repositories · arXiv:2310.02743Syntology official (archive's flag): 2 ran · 5 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 5 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 3 pointer-only (licence)
-
Adapting LLM Agents with Universal Feedback in Communication 1 Oct 2023 · 0 repositories · arXiv:2310.01444
-
Deep Reinforcement Learning for Autonomous Vehicle Intersection Navigation 30 Sep 2023 · 0 repositories · arXiv:2310.08595
-
Pairwise Proximal Policy Optimization: Harnessing Relative Feedback for LLM Alignment 30 Sep 2023 · 0 repositories · arXiv:2310.00212
-
CARLA: Adjusted common average referencing for cortico-cortical evoked potential data 29 Sep 2023 · 1 repository · arXiv:2310.00185
-
Cleanba: A Reproducible and Efficient Distributed Reinforcement Learning Platform 29 Sep 2023 · 1 repository · arXiv:2310.00036Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
Autonomous Driving using Spiking Neural Networks on Dynamic Vision Sensor Data: A Case Study of Traffic Light Change Detection 27 Sep 2023 · 1 repository · arXiv:2311.09225
-
Don't throw away your value model! Generating more preferable text with Value-Guided Monte-Carlo Tree Search decoding 26 Sep 2023 · 0 repositories · arXiv:2309.15028
-
Enhancing data efficiency in reinforcement learning: a novel imagination mechanism based on mesh information propagation 25 Sep 2023 · 2 repositories · arXiv:2309.14243
-
Boosting Offline Reinforcement Learning for Autonomous Driving with Hierarchical Latent Skills 24 Sep 2023 · 0 repositories · arXiv:2309.13614
-
Hierarchical Adaptive Value Estimation for Multi-modal Visual Reinforcement Learning 21 Sep 2023 · 1 repository
-
Learning Better with Less: Effective Augmentation for Sample-Efficient Visual Reinforcement Learning 21 Sep 2023 · 1 repository
-
Learning to Drive Anywhere 21 Sep 2023 · 0 repositories · arXiv:2309.12295
-
RRHF: Rank Responses to Align Language Models with Human Feedback 21 Sep 2023 · 1 repository
-
Adaptive Multi-Agent Deep Reinforcement Learning for Timely Healthcare Interventions 20 Sep 2023 · 0 repositories · arXiv:2309.10980
-
Privileged to Predicted: Towards Sensorimotor Reinforcement Learning for Urban Driving 18 Sep 2023 · 0 repositories · arXiv:2309.09756
-
Stabilizing RLHF through Advantage Model and Selective Rehearsal 18 Sep 2023 · 0 repositories · arXiv:2309.10202
-
Exploring the impact of low-rank adaptation on the performance, efficiency, and regularization of RLHF 16 Sep 2023 · 1 repository · arXiv:2309.09055Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 0 violated, 12 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 13 harvested samples)
-
What Matters to Enhance Traffic Rule Compliance of Imitation Learning for End-to-End Autonomous Driving 14 Sep 2023 · 0 repositories · arXiv:2309.07808
-
Deep Reinforcement Learning Enabled Joint Deployment and Beamforming in STAR-RIS Assisted Networks 7 Sep 2023 · 0 repositories · arXiv:2309.03520
-
SKoPe3D: A Synthetic Dataset for Vehicle Keypoint Perception in 3D from Traffic Monitoring Cameras 4 Sep 2023 · 0 repositories · arXiv:2309.01324
-
Efficient RLHF: Reducing the Memory Usage of PPO 1 Sep 2023 · 0 repositories · arXiv:2309.00754
-
A reinforcement learning based construction material supply strategy using robotic crane and computer vision for building reconstruction after an earthquake 30 Aug 2023 · 0 repositories · arXiv:2308.16280
-
On Offline Evaluation of 3D Object Detection for Autonomous Driving 24 Aug 2023 · 0 repositories · arXiv:2308.12779
-
Aligning Language Models with Offline Learning from Human Feedback 23 Aug 2023 · 2 repositories · arXiv:2308.12050
-
Pedestrian Environment Model for Automated Driving 17 Aug 2023 · 1 repository · arXiv:2308.09080
-
Robust Autonomous Vehicle Pursuit without Expert Steering Labels 16 Aug 2023 · 0 repositories · arXiv:2308.08380
-
A Deep Recurrent-Reinforcement Learning Method for Intelligent AutoScaling of Serverless Functions 11 Aug 2023 · 1 repository · arXiv:2308.05937
-
Proximal Policy Optimization Actual Combat: Manipulating Output Tokenizer Length 10 Aug 2023 · 0 repositories · arXiv:2308.05585
-
Exploring the Physical World Adversarial Robustness of Vehicle Detection 7 Aug 2023 · 0 repositories · arXiv:2308.03476
-
Cognitive TransFuser: Semantics-guided Transformer-based Sensor Fusion for Improved Waypoint Prediction 4 Aug 2023 · 1 repository · arXiv:2308.02126
-
Interpretable End-to-End Driving Model for Implicit Scene Understanding 2 Aug 2023 · 0 repositories · arXiv:2308.01180
-
Parallel Q-Learning: Scaling Off-policy Reinforcement Learning under Massively Parallel Simulation 24 Jul 2023 · 0 repositories · arXiv:2307.12983Syntology 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Trust-aware Safe Control for Autonomous Navigation: Estimation of System-to-human Trust for Trust-adaptive Control Barrier Functions 24 Jul 2023 · 3 repositories · arXiv:2307.12815
-
Llama 2: Open Foundation and Fine-Tuned Chat Models 18 Jul 2023 · 19 repositories · arXiv:2307.09288Syntology community repositories only · 33 ran (of which 7 constructed an object rather than computing a result; 22 with no instrument failure: 1 honoured, 1 violated, 20 with no contract checked; 11 where Syntology's instrument failed) · 19 unverified (of 52 harvested samples) · 20 pointer-only (licence)
-
Navigating Uncertainty: The Role of Short-Term Trajectory Prediction in Autonomous Vehicle Safety 11 Jul 2023 · 1 repository · arXiv:2307.05288
-
Secrets of RLHF in Large Language Models Part I: PPO 11 Jul 2023 · 1 repository · arXiv:2307.04964
-
Discovering Hierarchical Achievements in Reinforcement Learning via Contrastive Learning 7 Jul 2023 · 1 repository · arXiv:2307.03486
-
ContainerGym: A Real-World Reinforcement Learning Benchmark for Resource Allocation 6 Jul 2023 · 1 repository · arXiv:2307.02991
-
Navigation of micro-robot swarms for targeted delivery using reinforcement learning 30 Jun 2023 · 0 repositories · arXiv:2306.17598
-
Preference Ranking Optimization for Human Alignment 30 Jun 2023 · 1 repository · arXiv:2306.17492
-
Action and Trajectory Planning for Urban Autonomous Driving with Hierarchical Reinforcement Learning 28 Jun 2023 · 0 repositories · arXiv:2306.15968
-
Creating Valid Adversarial Examples of Malware 23 Jun 2023 · 1 repository · arXiv:2306.13587
-
Autonomous Driving with Deep Reinforcement Learning in CARLA Simulation 20 Jun 2023 · 0 repositories · arXiv:2306.11217
-
Learning to Generate Better Than Your LLM 20 Jun 2023 · 1 repository · arXiv:2306.11816
-
Safe, Efficient, Comfort, and Energy-saving Automated Driving through Roundabout Based on Deep Reinforcement Learning 20 Jun 2023 · 0 repositories · arXiv:2306.11465
-
A Study on Quantifying Sim2Real Image Gap in Autonomous Driving Simulations Using Lane Segmentation Attention Map Similarity 18 Jun 2023 · 0 repositories · arXiv:2306.10491
-
Coaching a Teachable Student 16 Jun 2023 · 1 repository · arXiv:2306.10014Syntology official: harvested, nothing ran · 0 ran · 4 unverified (of 4 harvested samples)
-
Hidden Biases of End-to-End Driving Models 13 Jun 2023 · 1 repository · arXiv:2306.07957Syntology official (archive's flag): 13 ran · 13 ran (of which 10 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 3 where Syntology's instrument failed) · 4 unverified (of 17 harvested samples)
-
RLtools: A Fast, Portable Deep Reinforcement Learning Library for Continuous Control 6 Jun 2023 · 1 repository · arXiv:2306.03530
-
Fine-Tuning Language Models with Advantage-Induced Policy Alignment 4 Jun 2023 · 1 repository · arXiv:2306.02231Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples) · 1 pointer-only (licence)
-
Deep Q-Learning versus Proximal Policy Optimization: Performance Comparison in a Material Sorting Task 2 Jun 2023 · 0 repositories · arXiv:2306.01451
-
ReLU to the Rescue: Improve Your On-Policy Actor-Critic with Positive Advantages 2 Jun 2023 · 1 repository · arXiv:2306.01460
-
Normalization Enhances Generalization in Visual Reinforcement Learning 1 Jun 2023 · 1 repository · arXiv:2306.00656Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Latent Exploration for Reinforcement Learning 31 May 2023 · 1 repository · arXiv:2305.20065
-
Resilience in Platoons of Cooperative Heterogeneous Vehicles: Self-organization Strategies and Provably-correct Design 27 May 2023 · 0 repositories · arXiv:2305.17443
-
Generating Synergistic Formulaic Alpha Collections via Reinforcement Learning 25 May 2023 · 1 repository · arXiv:2306.12964
-
Realistically distributing object placements in synthetic training data improves the performance of vision-based object detection models 24 May 2023 · 1 repository · arXiv:2305.14621
-
AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback 22 May 2023 · 2 repositories · arXiv:2305.14387
-
Learning Pedestrian Actions to Ensure Safe Autonomous Driving 22 May 2023 · 0 repositories · arXiv:2305.13051
-
Actor-Critic Methods using Physics-Informed Neural Networks: Control of a 1D PDE Model for Fluid-Cooled Battery Packs 18 May 2023 · 1 repository · arXiv:2305.10952
-
ReasonNet: End-to-End Driving with Temporal and Global Reasoning 17 May 2023 · 0 repositories · arXiv:2305.10507
-
SLiC-HF: Sequence Likelihood Calibration with Human Feedback 17 May 2023 · 0 repositories · arXiv:2305.10425
-
A Theoretical Analysis of Optimistic Proximal Policy Optimization in Linear Markov Decision Processes 15 May 2023 · 0 repositories · arXiv:2305.08841
-
Dynamically Conservative Self-Driving Planner for Long-Tail Cases 12 May 2023 · 0 repositories · arXiv:2305.07497
-
Think Twice before Driving: Towards Scalable Decoders for End-to-End Autonomous Driving 10 May 2023 · 1 repository · arXiv:2305.06242Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples)
-
Reducing the Cost of Cycle-Time Tuning for Real-World Policy Optimization 9 May 2023 · 1 repository · arXiv:2305.05760Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 2 pointer-only (licence)
-
Local Optimization Achieves Global Optimality in Multi-Agent Reinforcement Learning 8 May 2023 · 1 repository · arXiv:2305.04819Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
CARLA-BSP: a simulated dataset with pedestrians 29 Apr 2023 · 2 repositories · arXiv:2305.00204
-
Adversarial Policy Optimization in Deep Reinforcement Learning 27 Apr 2023 · 0 repositories · arXiv:2304.14533
-
Can Agents Run Relay Race with Strangers? Generalization of RL to Out-of-Distribution Trajectories 26 Apr 2023 · 0 repositories · arXiv:2304.13424
-
An End-to-End Vehicle Trajcetory Prediction Framework 19 Apr 2023 · 0 repositories · arXiv:2304.09764
-
Bridging RL Theory and Practice with the Effective Horizon 19 Apr 2023 · 1 repository · arXiv:2304.09853
-
Benchmarking the Physical-world Adversarial Robustness of Vehicle Detection 11 Apr 2023 · 0 repositories · arXiv:2304.05098
-
RRHF: Rank Responses to Align Language Models with Human Feedback without tears 11 Apr 2023 · 1 repository · arXiv:2304.05302Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
LANe: Lighting-Aware Neural Fields for Compositional Scene Synthesis 6 Apr 2023 · 0 repositories · arXiv:2304.03280
-
AutoRL Hyperparameter Landscapes 5 Apr 2023 · 1 repository · arXiv:2304.02396Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples)
-
PAC-Based Formal Verification for Out-of-Distribution Data Detection 4 Apr 2023 · 0 repositories · arXiv:2304.01592
-
Understanding Reinforcement Learning Algorithms: The Progress from Basic Q-learning to Proximal Policy Optimization 31 Mar 2023 · 0 repositories · arXiv:2304.00026
-
Specification-Guided Data Aggregation for Semantically Aware Imitation Learning 29 Mar 2023 · 0 repositories · arXiv:2303.17010
-
Model-Based Reinforcement Learning with Isolated Imaginations 27 Mar 2023 · 1 repository · arXiv:2303.14889