Methods › Reinforcement Learning › Policy Gradient Methods › PPO › Papers, page 3
Proximal Policy Optimization
PPO
Papers archive 2025-07-28
archive papers tagged: 949 · with a code link: 397 · where Syntology ran a sample: 139 (114 with a run with no instrument failure, 25 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (139 of 949 tagged: 114 with a run with no instrument failure, 25 where every run was a failure of Syntology's instrument)
Page 3 of 10: papers 201 to 300 of 949, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Benchmarking Deep Reinforcement Learning for Navigation in Denied Sensor Environments 18 Oct 2024 · 1 repository · arXiv:2410.14616
-
On the Influence of Shape, Texture and Color for Learning Semantic Segmentation 18 Oct 2024 · 0 repositories · arXiv:2410.14878
-
UniDrive: Towards Universal Driving Perception Across Camera Configurations 17 Oct 2024 · 1 repository · arXiv:2410.13864Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 8 harvested samples)
-
ASTM :Autonomous Smart Traffic Management System Using Artificial Intelligence CNN and LSTM 14 Oct 2024 · 0 repositories · arXiv:2410.10929
-
Improving the Language Understanding Capabilities of Large Language Models Using Reinforcement Learning 14 Oct 2024 · 0 repositories · arXiv:2410.11020
-
Taming Overconfidence in LLMs: Reward Calibration in RLHF 13 Oct 2024 · 1 repository · arXiv:2410.09724Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
Increasing the Difficulty of Automatically Generated Questions via Reinforcement Learning with Synthetic Preference 10 Oct 2024 · 0 repositories · arXiv:2410.08289
-
Adver-City: Open-Source Multi-Modal Dataset for Collaborative Perception Under Adverse Weather Conditions 8 Oct 2024 · 0 repositories · arXiv:2410.06380
-
Coevolving with the Other You: Fine-Tuning LLM with Sequential Cooperative Multi-Agent Reinforcement Learning 8 Oct 2024 · 1 repository · arXiv:2410.06101Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (of 11 harvested samples) · 1 pointer-only (licence)
-
Human-in-the-loop Reasoning For Traffic Sign Detection: Collaborative Approach Yolo With Video-llava 7 Oct 2024 · 0 repositories · arXiv:2410.05096
-
Mastering Chinese Chess AI (Xiangqi) Without Search 7 Oct 2024 · 0 repositories · arXiv:2410.04865
-
End-to-end Driving in High-Interaction Traffic Scenarios with Reinforcement Learning 3 Oct 2024 · 0 repositories · arXiv:2410.02253
-
VinePPO: Unlocking RL Potential For LLM Reasoning Through Refined Credit Assignment 2 Oct 2024 · 1 repository · arXiv:2410.01679Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 9 harvested samples)
-
The Perfect Blend: Redefining RLHF with Mixture of Judges 30 Sep 2024 · 0 repositories · arXiv:2409.20370
-
Spatial Reasoning and Planning for Deep Embodied Agents 28 Sep 2024 · 0 repositories · arXiv:2409.19479
-
Good Data Is All Imitation Learning Needs 26 Sep 2024 · 0 repositories · arXiv:2409.17605
-
Navigation in a simplified Urban Flow through Deep Reinforcement Learning 26 Sep 2024 · 0 repositories · arXiv:2409.17922
-
Dashing for the Golden Snitch: Multi-Drone Time-Optimal Motion Planning with Multi-Agent Reinforcement Learning 25 Sep 2024 · 1 repository · arXiv:2409.16720
-
Mitigating Covariate Shift in Imitation Learning for Autonomous Vehicles Using Latent Space Generative World Models 25 Sep 2024 · 0 repositories · arXiv:2409.16663
-
Safe Navigation for Robotic Digestive Endoscopy via Human Intervention-based Reinforcement Learning 24 Sep 2024 · 0 repositories · arXiv:2409.15688
-
Quantifying Context Bias in Domain Adaptation for Object Detection 23 Sep 2024 · 0 repositories · arXiv:2409.14679
-
LFP: Efficient and Accurate End-to-End Lane-Level Planning via Camera-LiDAR Fusion 21 Sep 2024 · 0 repositories · arXiv:2409.14170
-
METDrive: Multi-modal End-to-end Autonomous Driving with Temporal Guidance 19 Sep 2024 · 0 repositories · arXiv:2409.12667
-
Automating proton PBS treatment planning for head and neck cancers using policy gradient-based deep reinforcement learning 17 Sep 2024 · 0 repositories · arXiv:2409.11576
-
On-policy Actor-Critic Reinforcement Learning for Multi-UAV Exploration 17 Sep 2024 · 0 repositories · arXiv:2409.11058
-
Disentangling Uncertainty for Safe Social Navigation using Deep Reinforcement Learning 16 Sep 2024 · 0 repositories · arXiv:2409.10655
-
Risk-Aware Autonomous Driving with Linear Temporal Logic Specifications 15 Sep 2024 · 0 repositories · arXiv:2409.09769
-
Applying Action Masking and Curriculum Learning Techniques to Improve Data Efficiency and Overall Performance in Operational Technology Cyber Security using Reinforcement Learning 13 Sep 2024 · 0 repositories · arXiv:2409.10563
-
Module-wise Adaptive Adversarial Training for End-to-end Autonomous Driving 11 Sep 2024 · 0 repositories · arXiv:2409.07321
-
Multi-V2X: A Large Scale Multi-modal Multi-penetration-rate Dataset for Cooperative Perception 8 Sep 2024 · 1 repository · arXiv:2409.04980
-
QuantFactor REINFORCE: Mining Steady Formulaic Alpha Factors with Variance-bounded REINFORCE 8 Sep 2024 · 0 repositories · arXiv:2409.05144
-
Reinforcement Learning-enabled Satellite Constellation Reconfiguration and Retasking for Mission-Critical Applications 3 Sep 2024 · 0 repositories · arXiv:2409.02270
-
Development of Occupancy Prediction Algorithm for Underground Parking Lots 2 Sep 2024 · 0 repositories · arXiv:2409.00923
-
Enhancing Sample Efficiency and Exploration in Reinforcement Learning through the Integration of Diffusion Models and Proximal Policy Optimization 2 Sep 2024 · 1 repository · arXiv:2409.01427
-
MaFeRw: Query Rewriting with Multi-Aspect Feedbacks for Retrieval-Augmented Large Language Models 30 Aug 2024 · 1 repository · arXiv:2408.17072
-
An Extremely Data-efficient and Generative LLM-based Reinforcement Learning Agent for Recommenders 28 Aug 2024 · 0 repositories · arXiv:2408.16032
-
Comparison of Model Predictive Control and Proximal Policy Optimization for a 1-DOF Helicopter System 28 Aug 2024 · 0 repositories · arXiv:2408.15633
-
Instruct-SkillMix: A Powerful Pipeline for LLM Instruction Tuning 27 Aug 2024 · 1 repository · arXiv:2408.14774Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
Inverse-Q*: Token Level Reinforcement Learning for Aligning Large Language Models Without Preference Data 27 Aug 2024 · 0 repositories · arXiv:2408.14874
-
Pareto Inverse Reinforcement Learning for Diverse Expert Policy Generation 22 Aug 2024 · 0 repositories · arXiv:2408.12110
-
CARLA Drone: Monocular 3D Object Detection from a Different Perspective 21 Aug 2024 · 0 repositories · arXiv:2408.11958
-
Minor SFT loss for LLM fine-tune to increase performance and reduce model deviation 20 Aug 2024 · 0 repositories · arXiv:2408.10642
-
SynTraC: A Synthetic Dataset for Traffic Signal Control from Traffic Monitoring Cameras 18 Aug 2024 · 1 repository · arXiv:2408.09588
-
S-RAF: A Simulation-Based Robustness Assessment Framework for Responsible Autonomous Driving 16 Aug 2024 · 1 repository · arXiv:2408.08584
-
Enhancing Autonomous Vehicle Perception in Adverse Weather through Image Augmentation during Semantic Segmentation Training 14 Aug 2024 · 1 repository · arXiv:2408.07239
-
Kolmogorov-Arnold Network for Online Reinforcement Learning 9 Aug 2024 · 1 repository · arXiv:2408.04841
-
Cross-cultural analysis of pedestrian group behaviour influence on crossing decisions in interactions with autonomous vehicles 6 Aug 2024 · 0 repositories · arXiv:2408.03003
-
Research on Autonomous Driving Decision-making Strategies based Deep Reinforcement Learning 6 Aug 2024 · 0 repositories · arXiv:2408.03084
-
MSMA: Multi-agent Trajectory Prediction in Connected and Autonomous Vehicle Environment with Multi-source Data Integration 31 Jul 2024 · 1 repository · arXiv:2407.21310
-
Appraisal-Guided Proximal Policy Optimization: Modeling Psychological Disorders in Dynamic Grid World 29 Jul 2024 · 0 repositories · arXiv:2407.20383
-
SAPG: Split and Aggregate Policy Gradients 29 Jul 2024 · 0 repositories · arXiv:2407.20230
-
Towards Aligning Language Models with Textual Feedback 24 Jul 2024 · 1 repository · arXiv:2407.16970Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
In Search for Architectures and Loss Functions in Multi-Objective Reinforcement Learning 23 Jul 2024 · 0 repositories · arXiv:2407.16807
-
Enhancing Hardware Fault Tolerance in Machines with Reinforcement Learning Policy Gradient Algorithms 21 Jul 2024 · 0 repositories · arXiv:2407.15283
-
Understanding Reinforcement Learning-Based Fine-Tuning of Diffusion Models: A Tutorial and Review 18 Jul 2024 · 1 repository · arXiv:2407.13734
-
Perception Helps Planning: Facilitating Multi-Stage Lane-Level Integration via Double-Edge Structures 16 Jul 2024 · 0 repositories · arXiv:2407.11644
-
DINO Pre-training for Vision-based End-to-end Autonomous Driving 15 Jul 2024 · 0 repositories · arXiv:2407.10803
-
Digital twins to alleviate the need for real field data in vision-based vehicle speed detection systems 11 Jul 2024 · 0 repositories · arXiv:2407.08380
-
Real-time system optimal traffic routing under uncertainties -- Can physics models boost reinforcement learning? 10 Jul 2024 · 0 repositories · arXiv:2407.07364
-
Towards Human-Like Driving: Active Inference in Autonomous Vehicle Control 10 Jul 2024 · 0 repositories · arXiv:2407.07684
-
Exploring the Causality of End-to-End Autonomous Driving 9 Jul 2024 · 1 repository · arXiv:2407.06546
-
Exposing Privacy Gaps: Membership Inference Attack on Preference Data for LLM Alignment 8 Jul 2024 · 0 repositories · arXiv:2407.06443
-
Simplifying Deep Temporal Difference Learning 5 Jul 2024 · 1 repository · arXiv:2407.04811Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Efficient Fusion and Task Guided Embedding for End-to-end Autonomous Driving 3 Jul 2024 · 0 repositories · arXiv:2407.02878
-
PPO-based Dynamic Control of Uncertain Floating Platforms in the Zero-G Environment 3 Jul 2024 · 0 repositories · arXiv:2407.03224
-
A Deep Reinforcement Learning Approach to Battery Management in Dairy Farming via Proximal Policy Optimization 1 Jul 2024 · 0 repositories · arXiv:2407.01653
-
Acceleration method for generating perception failure scenarios based on editing Markov process 1 Jul 2024 · 0 repositories · arXiv:2407.00980
-
Deep Reinforcement Learning for Adverse Garage Scenario Generation 1 Jul 2024 · 0 repositories · arXiv:2407.01333
-
Deep Reinforcement Learning Strategies in Finance: Insights into Asset Holding, Trading Behavior, and Purchase Diversity 29 Jun 2024 · 0 repositories · arXiv:2407.09557
-
Decoding-Time Language Model Alignment with Multiple Objectives 27 Jun 2024 · 1 repository · arXiv:2406.18853Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample)
-
End-to-End Autonomous Driving without Costly Modularization and 3D Manual Annotation 25 Jun 2024 · 0 repositories · arXiv:2406.17680
-
Performance Comparison of Deep RL Algorithms for Mixed Traffic Cooperative Lane-Changing 25 Jun 2024 · 0 repositories · arXiv:2407.02521
-
RaCIL: Ray Tracing based Multi-UAV Obstacle Avoidance through Composite Imitation Learning 24 Jun 2024 · 0 repositories · arXiv:2407.02520
-
Multistep Criticality Search and Power Shaping in Microreactors with Reinforcement Learning 22 Jun 2024 · 0 repositories · arXiv:2406.15931
-
SeRTS: Self-Rewarding Tree Search for Biomedical Retrieval-Augmented Generation 17 Jun 2024 · 0 repositories · arXiv:2406.11258
-
P-TA: Using Proximal Policy Optimization to Enhance Tabular Data Augmentation via Large Language Models 17 Jun 2024 · 0 repositories · arXiv:2406.11391
-
Planning with Adaptive World Models for Autonomous Driving 15 Jun 2024 · 0 repositories · arXiv:2406.10714
-
CarLLaVA: Vision language models for camera-only closed-loop driving 14 Jun 2024 · 1 repository · arXiv:2406.10165
-
Hadamard Representations: Augmenting Hyperbolic Tangents in RL 13 Jun 2024 · 0 repositories · arXiv:2406.09079
-
Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback 13 Jun 2024 · 2 repositories · arXiv:2406.09279Syntology official (archive's flag): 15 ran · 15 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 7 where Syntology's instrument failed) · 4 unverified (of 19 harvested samples)
-
Optimizing Deep Reinforcement Learning for Adaptive Robotic Arm Control 12 Jun 2024 · 0 repositories · arXiv:2407.02503
-
PRIBOOT: A New Data-Driven Expert for Improved Driving Simulations 12 Jun 2024 · 1 repository · arXiv:2406.08421
-
Adaptive Opponent Policy Detection in Multi-Agent MDPs: Real-Time Strategy Switch Identification Using Running Error Estimation 10 Jun 2024 · 0 repositories · arXiv:2406.06500
-
Flow of Reasoning:Training LLMs for Divergent Problem Solving with Minimal Examples 9 Jun 2024 · 1 repository · arXiv:2406.05673Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 15 harvested samples)
-
Online Policy Distillation with Decision-Attention 8 Jun 2024 · 0 repositories · arXiv:2406.05488
-
Bench2Drive: Towards Multi-Ability Benchmarking of Closed-Loop End-To-End Autonomous Driving 6 Jun 2024 · 4 repositories · arXiv:2406.03877Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 3 where Syntology's instrument failed) · 4 unverified (of 17 harvested samples) · 17 pointer-only (licence)
-
HackAtari: Atari Learning Environments for Robust and Continual Reinforcement Learning 6 Jun 2024 · 1 repository · arXiv:2406.03997Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Transductive Off-policy Proximal Policy Optimization 6 Jun 2024 · 0 repositories · arXiv:2406.03894
-
Prompt-based Visual Alignment for Zero-shot Policy Transfer 5 Jun 2024 · 0 repositories · arXiv:2406.03250
-
Aligning Large Language Models via Fine-grained Supervision 4 Jun 2024 · 0 repositories · arXiv:2406.02756
-
LanEvil: Benchmarking the Robustness of Lane Detection to Environmental Illusions 3 Jun 2024 · 0 repositories · arXiv:2406.00934
-
Validity Learning on Failures: Mitigating the Distribution Shift in Autonomous Vehicle Planning 3 Jun 2024 · 0 repositories · arXiv:2406.01544
-
Self-Improving Robust Preference Optimization 3 Jun 2024 · 0 repositories · arXiv:2406.01660
-
Pedestrian intention prediction in Adverse Weather Conditions with Spiking Neural Networks and Dynamic Vision Sensors 1 Jun 2024 · 0 repositories · arXiv:2406.00473
-
Aquatic Navigation: A Challenging Benchmark for Deep Reinforcement Learning 30 May 2024 · 1 repository · arXiv:2405.20534
-
Intrinsic Dynamics-Driven Generalizable Scene Representations for Vision-Oriented Decision-Making Applications 30 May 2024 · 1 repository · arXiv:2405.19736
-
Linear Function Approximation as a Computationally Efficient Method to Solve Classical Reinforcement Learning Challenges 27 May 2024 · 0 repositories · arXiv:2405.20350
-
SCaRL- A Synthetic Multi-Modal Dataset for Autonomous Driving 27 May 2024 · 0 repositories · arXiv:2405.17030
-
Symmetric Reinforcement Learning Loss for Robust Learning on Diverse Tasks and Model Scales 27 May 2024 · 1 repository · arXiv:2405.17618
-
Rewarded Region Replay (R3) for Policy Learning with Discrete Action Space 26 May 2024 · 1 repository · arXiv:2405.16383