Methods › Reinforcement Learning › Policy Gradient Methods › PPO › Papers, page 4
Proximal Policy Optimization
PPO
Papers archive 2025-07-28
archive papers tagged: 949 · with a code link: 397 · where Syntology ran a sample: 139 (114 with a run with no instrument failure, 25 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (139 of 949 tagged: 114 with a run with no instrument failure, 25 where every run was a failure of Syntology's instrument)
Page 4 of 10: papers 301 to 400 of 949, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Continuously Learning, Adapting, and Improving: A Dual-Process Approach to Autonomous Driving 24 May 2024 · 1 repository · arXiv:2405.15324
-
Extracting Heuristics from Large Language Models for Reward Shaping in Reinforcement Learning 24 May 2024 · 0 repositories · arXiv:2405.15194
-
AGILE: A Novel Reinforcement Learning Framework of LLM Agents 23 May 2024 · 1 repository · arXiv:2405.14751Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 1 pointer-only (licence)
-
ChatScene: Knowledge-Enabled Safety-Critical Scenario Generation for Autonomous Vehicles 22 May 2024 · 1 repository · arXiv:2405.14062Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample)
-
MaskFuser: Masked Fusion of Joint Multi-Modal Tokenization for End-to-End Autonomous Driving 13 May 2024 · 0 repositories · arXiv:2405.07573
-
Value Augmented Sampling for Language Model Alignment and Personalization 10 May 2024 · 1 repository · arXiv:2405.06639Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Proximal Policy Optimization with Adaptive Exploration 7 May 2024 · 1 repository · arXiv:2405.04664
-
Guidance Design for Escape Flight Vehicle Using Evolution Strategy Enhanced Deep Reinforcement Learning 4 May 2024 · 0 repositories · arXiv:2405.03711
-
D2PO: Discriminator-Guided DPO with Response Evaluation Models 2 May 2024 · 1 repository · arXiv:2405.01511Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 13 harvested samples) · 13 pointer-only (licence)
-
No Representation, No Trust: Connecting Representation, Collapse, and Trust Issues in PPO 1 May 2024 · 1 repository · arXiv:2405.00662Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 8 harvested samples) · 1 pointer-only (licence)
-
Guiding Attention in End-to-End Driving Models 30 Apr 2024 · 1 repository · arXiv:2405.00242
-
DPO Meets PPO: Reinforced Token Optimization for RLHF 29 Apr 2024 · 1 repository · arXiv:2404.18922Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples) · 1 pointer-only (licence)
-
3D Extended Object Tracking by Fusing Roadside Sparse Radar Point Clouds and Pixel Keypoints 27 Apr 2024 · 2 repositories · arXiv:2404.17903
-
REBEL: Reinforcement Learning via Regressing Relative Rewards 25 Apr 2024 · 3 repositories · arXiv:2404.16767Syntology official (archive's flag): 16 ran · 16 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 1 honoured, 0 violated, 11 with no contract checked; 4 where Syntology's instrument failed) · 4 unverified (of 20 harvested samples) · 6 pointer-only (licence)
-
ContextualFusion: Context-Based Multi-Sensor Fusion for 3D Object Detection in Adverse Operating Conditions 23 Apr 2024 · 0 repositories · arXiv:2404.14780
-
Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study 16 Apr 2024 · 1 repository · arXiv:2404.10719
-
SEVD: Synthetic Event-based Vision Dataset for Ego and Fixed Traffic Perception 12 Apr 2024 · 1 repository · arXiv:2404.10540
-
Enhanced Cooperative Perception for Autonomous Vehicles Using Imperfect Communication 10 Apr 2024 · 0 repositories · arXiv:2404.08013
-
Synergy of Large Language Model and Model Driven Engineering for Automated Development of Centralized Vehicular Systems 8 Apr 2024 · 0 repositories · arXiv:2404.05508
-
Prompting Multi-Modal Tokens to Enhance End-to-End Autonomous Driving Imitation Learning with LLMs 7 Apr 2024 · 0 repositories · arXiv:2404.04869
-
A proximal policy optimization based intelligent home solar management 5 Apr 2024 · 0 repositories · arXiv:2404.03888
-
AGL-NET: Aerial-Ground Cross-Modal Global Localization with Varying Scales 4 Apr 2024 · 1 repository · arXiv:2404.03187
-
Addressing Loss of Plasticity and Catastrophic Forgetting in Continual Learning 31 Mar 2024 · 1 repository · arXiv:2404.00781Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample)
-
Human-compatible driving partners through data-regularized self-play reinforcement learning 28 Mar 2024 · 1 repository · arXiv:2403.19648
-
Scenario-Based Curriculum Generation for Multi-Agent Autonomous Driving 26 Mar 2024 · 1 repository · arXiv:2403.17805
-
DriveCoT: Integrating Chain-of-Thought Reasoning with End-to-End Driving 25 Mar 2024 · 0 repositories · arXiv:2403.16996
-
Policy Mirror Descent with Lookahead 21 Mar 2024 · 1 repository · arXiv:2403.14156
-
Equivariant Ensembles and Regularization for Reinforcement Learning in Map-based Path Planning 19 Mar 2024 · 1 repository · arXiv:2403.12856
-
JaxUED: A simple and useable UED library in Jax 19 Mar 2024 · 1 repository · arXiv:2403.13091
-
M2DA: Multi-Modal Fusion Transformer Incorporating Driver Attention for Autonomous Driving 19 Mar 2024 · 0 repositories · arXiv:2403.12552
-
Driving Style Alignment for LLM-powered Driver Agent 17 Mar 2024 · 2 repositories · arXiv:2403.11368
-
Are you a robot? Detecting Autonomous Vehicles from Behavior Analysis 14 Mar 2024 · 0 repositories · arXiv:2403.09571
-
Right Place, Right Time! Dynamizing Topological Graphs for Embodied Navigation 14 Mar 2024 · 0 repositories · arXiv:2403.09905
-
Improving Reinforcement Learning from Human Feedback Using Contrastive Rewards 12 Mar 2024 · 0 repositories · arXiv:2403.07708
-
Tractable Joint Prediction and Planning over Discrete Behavior Modes for Urban Driving 12 Mar 2024 · 0 repositories · arXiv:2403.07232
-
Risk-Sensitive RL with Optimized Certainty Equivalents via Reduction to Standard RL 10 Mar 2024 · 0 repositories · arXiv:2403.06323
-
Teaching Large Language Models to Reason with Reinforcement Learning 7 Mar 2024 · 0 repositories · arXiv:2403.04642
-
COMMIT: Certifying Robustness of Multi-Sensor Fusion Systems against Semantic Attacks 4 Mar 2024 · 0 repositories · arXiv:2403.02329
-
Snapshot Reinforcement Learning: Leveraging Prior Trajectories for Efficiency 1 Mar 2024 · 1 repository · arXiv:2403.00673
-
Craftax: A Lightning-Fast Benchmark for Open-Ended Reinforcement Learning 26 Feb 2024 · 1 repository · arXiv:2402.16801Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 14 unverified (of 16 harvested samples)
-
Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs 22 Feb 2024 · 0 repositories · arXiv:2402.14740
-
Distributed Radiance Fields for Edge Video Compression and Metaverse Integration in Autonomous Driving 22 Feb 2024 · 0 repositories · arXiv:2402.14642
-
Hybrid Reasoning Based on Large Language Models for Autonomous Car Driving 21 Feb 2024 · 1 repository · arXiv:2402.13602
-
VADv2: End-to-End Vectorized Autonomous Driving via Probabilistic Planning 20 Feb 2024 · 1 repository · arXiv:2402.13243Syntology official (archive's flag): 1 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Surpassing legacy approaches to PWR core reload optimization with single-objective Reinforcement learning 16 Feb 2024 · 0 repositories · arXiv:2402.11040
-
RS-DPO: A Hybrid Rejection Sampling and Direct Preference Optimization Method for Alignment of Large Language Models 15 Feb 2024 · 0 repositories · arXiv:2402.10038
-
Reducing Texture Bias of Deep Neural Networks via Edge Enhancing Diffusion 14 Feb 2024 · 1 repository · arXiv:2402.09530
-
SemTra: A Semantic Skill Translator for Cross-Domain Zero-Shot Policy Adaptation 12 Feb 2024 · 0 repositories · arXiv:2402.07418
-
Solving Deep Reinforcement Learning Tasks with Evolution Strategies and Linear Policy Networks 10 Feb 2024 · 1 repository · arXiv:2402.06912
-
Entropy-Regularized Token-Level Policy Optimization for Language Agent Reinforcement 9 Feb 2024 · 1 repository · arXiv:2402.06700Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 1 violated, 2 with no contract checked; 3 where Syntology's instrument failed) · 3 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
Attention-Enhanced Prioritized Proximal Policy Optimization for Adaptive Edge Caching 8 Feb 2024 · 0 repositories · arXiv:2402.14576
-
Convergence for Natural Policy Gradient on Infinite-State Queueing MDPs 7 Feb 2024 · 0 repositories · arXiv:2402.05274
-
Tuning the feedback controller gains is a simple way to improve autonomous driving performance 7 Feb 2024 · 0 repositories · arXiv:2402.05064
-
Averaging n-step Returns Reduces Variance in Reinforcement Learning 6 Feb 2024 · 0 repositories · arXiv:2402.03903Syntology 5 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 8 harvested samples) · 6 pointer-only (licence)
-
Learning to Generate Explainable Stock Predictions using Self-Reflective Large Language Models 6 Feb 2024 · 1 repository · arXiv:2402.03659Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
OASim: an Open and Adaptive Simulator based on Neural Rendering for Autonomous Driving 6 Feb 2024 · 1 repository · arXiv:2402.03830
-
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models 5 Feb 2024 · 5 repositories · arXiv:2402.03300Syntology official (archive's flag): 8 ran · 14 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 4 where Syntology's instrument failed) · 10 unverified (of 24 harvested samples) · 3 pointer-only (licence)
-
BRAIn: Bayesian Reward-conditioned Amortized Inference for natural language generation from feedback 4 Feb 2024 · 0 repositories · arXiv:2402.02479
-
Hybrid-Prediction Integrated Planning for Autonomous Driving 4 Feb 2024 · 0 repositories · arXiv:2402.02426
-
Parametric-Task MAP-Elites 2 Feb 2024 · 0 repositories · arXiv:2402.01275
-
CARFF: Conditional Auto-encoded Radiance Field for 3D Scene Forecasting 31 Jan 2024 · 0 repositories · arXiv:2401.18075
-
Simple Policy Optimization 29 Jan 2024 · 1 repository · arXiv:2401.16025
-
True Knowledge Comes from Practice: Aligning LLMs with Embodied Environments via Reinforcement Learning 25 Jan 2024 · 1 repository · arXiv:2401.14151Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Linear Alignment: A Closed-form Solution for Aligning Human Preferences without Tuning and Feedback 21 Jan 2024 · 1 repository · arXiv:2401.11458Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
360ORB-SLAM: A Visual SLAM System for Panoramic Images with Depth Completion Network 19 Jan 2024 · 0 repositories · arXiv:2401.10560
-
LangProp: A code optimization framework using Large Language Models applied to driving 18 Jan 2024 · 1 repository · arXiv:2401.10314Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
ReFT: Reasoning with Reinforced Fine-Tuning 17 Jan 2024 · 1 repository · arXiv:2401.08967Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 9 unverified (of 14 harvested samples) · 13 pointer-only (licence)
-
CPPO: Continual Learning for Reinforcement Learning with Human Feedback 16 Jan 2024 · 0 repositories
-
Sum Throughput Maximization in Multi-BD Symbiotic Radio NOMA Network Assisted by Active-STAR-RIS 16 Jan 2024 · 0 repositories · arXiv:2401.08301
-
Beyond Sparse Rewards: Enhancing Reinforcement Learning with Language Model Critique in Text Generation 14 Jan 2024 · 0 repositories · arXiv:2401.07382
-
Aquarium: A Comprehensive Framework for Exploring Predator-Prey Dynamics through Multi-Agent Reinforcement Learning Algorithms 13 Jan 2024 · 1 repository · arXiv:2401.07056
-
MAPO: Advancing Multilingual Reasoning through Multilingual Alignment-as-Preference Optimization 12 Jan 2024 · 1 repository · arXiv:2401.06838Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
Autonomous Navigation of Tractor-Trailer Vehicles through Roundabout Intersections 10 Jan 2024 · 0 repositories · arXiv:2401.04980
-
Feedback-Guided Autonomous Driving 1 Jan 2024 · 0 repositories
-
Beyond PID Controllers: PPO with Neuralized PID Policy for Proton Beam Intensity Control in Mu2e 28 Dec 2023 · 0 repositories · arXiv:2312.17372
-
Preference as Reward, Maximum Preference Optimization with Importance Sampling 27 Dec 2023 · 0 repositories · arXiv:2312.16430
-
Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss 27 Dec 2023 · 1 repository · arXiv:2312.16682Syntology 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 8 harvested samples) · 1 pointer-only (licence)
-
Agent based modelling for continuously varying supply chains 24 Dec 2023 · 0 repositories · arXiv:2312.15502
-
DriveLM: Driving with Graph Visual Question Answering 21 Dec 2023 · 3 repositories · arXiv:2312.14150
-
Realistic Rainy Weather Simulation for LiDARs in CARLA Simulator 20 Dec 2023 · 1 repository · arXiv:2312.12772
-
Colored Noise in PPO: Improved Exploration and Performance through Correlated Action Sampling 18 Dec 2023 · 1 repository · arXiv:2312.11091
-
Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint 18 Dec 2023 · 3 repositories · arXiv:2312.11456
-
DriveMLM: Aligning Multi-Modal Large Language Models with Behavioral Planning States for Autonomous Driving 14 Dec 2023 · 2 repositories · arXiv:2312.09245Syntology community repositories only · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Gradient Informed Proximal Policy Optimization 14 Dec 2023 · 1 repository · arXiv:2312.08710Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations 14 Dec 2023 · 3 repositories · arXiv:2312.08935
-
Challenges of YOLO Series for Object Detection in Extremely Heavy Rain: CALRA Simulator based Synthetic Evaluation Dataset 13 Dec 2023 · 0 repositories · arXiv:2312.07976
-
The Effective Horizon Explains Deep RL Performance in Stochastic Environments 13 Dec 2023 · 1 repository · arXiv:2312.08369
-
A dynamical clipping approach with task feedback for Proximal Policy Optimization 12 Dec 2023 · 0 repositories · arXiv:2312.07624
-
SkyScenes: A Synthetic Dataset for Aerial Scene Understanding 11 Dec 2023 · 0 repositories · arXiv:2312.06719
-
Graph-based Prediction and Planning Policy Network (GP3Net) for scalable self-driving in dynamic environments using Deep Reinforcement Learning 10 Dec 2023 · 0 repositories · arXiv:2312.05784
-
Can language agents be alternatives to PPO? A Preliminary Empirical Study On OpenAI Gym 6 Dec 2023 · 1 repository · arXiv:2312.03290
-
A Reliable Representation with Bidirectional Transition Model for Visual Reinforcement Learning Generalization 4 Dec 2023 · 0 repositories · arXiv:2312.01915
-
Data-efficient Deep Reinforcement Learning for Vehicle Trajectory Control 30 Nov 2023 · 0 repositories · arXiv:2311.18393
-
EpiTESTER: Testing Autonomous Vehicles with Epigenetic Algorithm and Attention Mechanism 30 Nov 2023 · 1 repository · arXiv:2312.00207
-
Safe Reinforcement Learning in a Simulated Robotic Arm 28 Nov 2023 · 0 repositories · arXiv:2312.09468
-
Automated Lane Merging via Game Theory and Branch Model Predictive Control 25 Nov 2023 · 1 repository · arXiv:2311.14916
-
Attacking Motion Planners Using Adversarial Perception Errors 21 Nov 2023 · 0 repositories · arXiv:2311.12722
-
Nav-Q: Quantum Deep Reinforcement Learning for Collision-Free Navigation of Self-Driving Cars 20 Nov 2023 · 1 repository · arXiv:2311.12875
-
Bridging Data-Driven and Knowledge-Driven Approaches for Safety-Critical Scenario Generation in Automated Vehicle Validation 18 Nov 2023 · 0 repositories · arXiv:2311.10937
-
Automatic Generation of Scenarios for System-level Simulation-based Verification of Autonomous Driving Systems 16 Nov 2023 · 0 repositories · arXiv:2311.09784