Methods › Reinforcement Learning › Policy Gradient Methods › PPO › Papers where code ran, page 1
Proximal Policy Optimization
PPO
Papers archive 2025-07-28
archive papers tagged: 949 · with a code link: 397 · where Syntology ran a sample: 139 (114 with a run with no instrument failure, 25 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (139 of 949 tagged: 114 with a run with no instrument failure, 25 where every run was a failure of Syntology's instrument)
Syntology We ran code from the paper's repository; we did not isolate this method inside it.
Page 1 of 2: papers 1 to 100 of the 139 tagged papers where Syntology ran at least one harvested sample (114 with a run with no instrument failure, 25 where every run was a failure of Syntology's instrument), newest first by the archive's date (ties by arXiv id). This is a filter on Syntology's record ordered by date only, not a ranking; a run is not a correctness claim. A paper missing from this list is not a recorded non-run: it may have no arXiv id, no harvested code, or only samples that have not run yet.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization 17 Jun 2025 · 1 repository · arXiv:2506.14574Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
AutoVLA: A Vision-Language-Action Model for End-to-End Autonomous Driving with Adaptive Reasoning and Reinforcement Fine-Tuning 16 Jun 2025 · 1 repository · arXiv:2506.13757Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
The Surprising Effectiveness of Negative Reinforcement in LLM Reasoning 2 Jun 2025 · 1 repository · arXiv:2506.01347Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning 30 May 2025 · 1 repository · arXiv:2505.24298Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 13 harvested samples)
-
Segment Policy Optimization: Effective Segment-Level Credit Assignment in RL for Large Language Models 29 May 2025 · 1 repository · arXiv:2505.23564Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 9 harvested samples)
-
Trajectory Bellman Residual Minimization: A Simple Value-Based Method for LLM Reasoning 21 May 2025 · 0 repositories · arXiv:2505.15311Syntology 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Reinforcement Learning Finetunes Small Subnetworks in Large Language Models 16 May 2025 · 0 repositories · arXiv:2505.11711Syntology 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
CaRL: Learning Scalable Planning Policies with Simple Rewards 24 Apr 2025 · 1 repository · arXiv:2504.17838Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce 15 Apr 2025 · 1 repository · arXiv:2504.11343Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Concise Reasoning via Reinforcement Learning 7 Apr 2025 · 1 repository · arXiv:2504.05185Syntology official (archive's flag): 11 ran · 12 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 1 violated, 10 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 14 harvested samples) · 13 pointer-only (licence)
-
End-to-End Driving with Online Trajectory Evaluation via BEV World Model 2 Apr 2025 · 1 repository · arXiv:2504.01941Syntology official (archive's flag): 7 ran · 7 ran (of which 7 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified; every one of the 7 samples that ran constructed an object rather than computing a result (of 11 harvested samples)
-
Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model 31 Mar 2025 · 1 repository · arXiv:2503.24290Syntology 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 4 harvested samples)
-
UniOcc: A Unified Benchmark for Occupancy Forecasting and Prediction in Autonomous Driving 31 Mar 2025 · 1 repository · arXiv:2503.24381Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples)
-
SimLingo: Vision-Only Closed-Loop Autonomous Driving with Language-Action Alignment 12 Mar 2025 · 1 repository · arXiv:2503.09594Syntology 11 ran (of which 8 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 3 where Syntology's instrument failed) · 3 unverified (of 14 harvested samples)
-
Reevaluating Policy Gradient Methods for Imperfect-Information Games 13 Feb 2025 · 2 repositories · arXiv:2502.08938Syntology official (archive's flag): 1 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
On the Emergence of Thinking in LLMs I: Searching for the Right Intuition 10 Feb 2025 · 4 repositories · arXiv:2502.06773Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 13 harvested samples) · 9 pointer-only (licence)
-
LeapVAD: A Leap in Autonomous Driving via Cognitive Perception and Dual-Process Thinking 14 Jan 2025 · 1 repository · arXiv:2501.08168Syntology 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 1 pointer-only (licence)
-
Efficient Policy Adaptation with Contrastive Prompt Ensemble for Embodied Agents 16 Dec 2024 · 0 repositories · arXiv:2412.11484Syntology 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample) · 1 pointer-only (licence)
-
WiseAD: Knowledge Augmented End-to-End Autonomous Driving with Vision-Language Model 13 Dec 2024 · 1 repository · arXiv:2412.09951Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 1 honoured, 0 violated, 2 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Hidden Biases of End-to-End Driving Datasets 12 Dec 2024 · 1 repository · arXiv:2412.09602Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 2 honoured, 1 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning 10 Dec 2024 · 1 repository · arXiv:2412.07165Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
CALE: Continuous Arcade Learning Environment 31 Oct 2024 · 1 repository · arXiv:2410.23810Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Fast Best-of-N Decoding via Speculative Rejection 26 Oct 2024 · 1 repository · arXiv:2410.20290Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
UniDrive: Towards Universal Driving Perception Across Camera Configurations 17 Oct 2024 · 1 repository · arXiv:2410.13864Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 8 harvested samples)
-
Coevolving with the Other You: Fine-Tuning LLM with Sequential Cooperative Multi-Agent Reinforcement Learning 8 Oct 2024 · 1 repository · arXiv:2410.06101Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (of 11 harvested samples) · 1 pointer-only (licence)
-
VinePPO: Unlocking RL Potential For LLM Reasoning Through Refined Credit Assignment 2 Oct 2024 · 1 repository · arXiv:2410.01679Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 9 harvested samples)
-
Instruct-SkillMix: A Powerful Pipeline for LLM Instruction Tuning 27 Aug 2024 · 1 repository · arXiv:2408.14774Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
Simplifying Deep Temporal Difference Learning 5 Jul 2024 · 1 repository · arXiv:2407.04811Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Decoding-Time Language Model Alignment with Multiple Objectives 27 Jun 2024 · 1 repository · arXiv:2406.18853Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample)
-
Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback 13 Jun 2024 · 2 repositories · arXiv:2406.09279Syntology official (archive's flag): 15 ran · 15 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 7 where Syntology's instrument failed) · 4 unverified (of 19 harvested samples)
-
Flow of Reasoning:Training LLMs for Divergent Problem Solving with Minimal Examples 9 Jun 2024 · 1 repository · arXiv:2406.05673Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 15 harvested samples)
-
Bench2Drive: Towards Multi-Ability Benchmarking of Closed-Loop End-To-End Autonomous Driving 6 Jun 2024 · 4 repositories · arXiv:2406.03877Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 3 where Syntology's instrument failed) · 4 unverified (of 17 harvested samples) · 17 pointer-only (licence)
-
HackAtari: Atari Learning Environments for Robust and Continual Reinforcement Learning 6 Jun 2024 · 1 repository · arXiv:2406.03997Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
AGILE: A Novel Reinforcement Learning Framework of LLM Agents 23 May 2024 · 1 repository · arXiv:2405.14751Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 1 pointer-only (licence)
-
ChatScene: Knowledge-Enabled Safety-Critical Scenario Generation for Autonomous Vehicles 22 May 2024 · 1 repository · arXiv:2405.14062Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample)
-
Value Augmented Sampling for Language Model Alignment and Personalization 10 May 2024 · 1 repository · arXiv:2405.06639Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
D2PO: Discriminator-Guided DPO with Response Evaluation Models 2 May 2024 · 1 repository · arXiv:2405.01511Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 13 harvested samples) · 13 pointer-only (licence)
-
No Representation, No Trust: Connecting Representation, Collapse, and Trust Issues in PPO 1 May 2024 · 1 repository · arXiv:2405.00662Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 8 harvested samples) · 1 pointer-only (licence)
-
DPO Meets PPO: Reinforced Token Optimization for RLHF 29 Apr 2024 · 1 repository · arXiv:2404.18922Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples) · 1 pointer-only (licence)
-
REBEL: Reinforcement Learning via Regressing Relative Rewards 25 Apr 2024 · 3 repositories · arXiv:2404.16767Syntology official (archive's flag): 16 ran · 16 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 1 honoured, 0 violated, 11 with no contract checked; 4 where Syntology's instrument failed) · 4 unverified (of 20 harvested samples) · 6 pointer-only (licence)
-
Addressing Loss of Plasticity and Catastrophic Forgetting in Continual Learning 31 Mar 2024 · 1 repository · arXiv:2404.00781Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample)
-
Craftax: A Lightning-Fast Benchmark for Open-Ended Reinforcement Learning 26 Feb 2024 · 1 repository · arXiv:2402.16801Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 14 unverified (of 16 harvested samples)
-
VADv2: End-to-End Vectorized Autonomous Driving via Probabilistic Planning 20 Feb 2024 · 1 repository · arXiv:2402.13243Syntology official (archive's flag): 1 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Entropy-Regularized Token-Level Policy Optimization for Language Agent Reinforcement 9 Feb 2024 · 1 repository · arXiv:2402.06700Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 1 violated, 2 with no contract checked; 3 where Syntology's instrument failed) · 3 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
Learning to Generate Explainable Stock Predictions using Self-Reflective Large Language Models 6 Feb 2024 · 1 repository · arXiv:2402.03659Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Averaging n-step Returns Reduces Variance in Reinforcement Learning 6 Feb 2024 · 0 repositories · arXiv:2402.03903Syntology 5 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 8 harvested samples) · 6 pointer-only (licence)
-
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models 5 Feb 2024 · 5 repositories · arXiv:2402.03300Syntology official (archive's flag): 8 ran · 14 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 4 where Syntology's instrument failed) · 10 unverified (of 24 harvested samples) · 3 pointer-only (licence)
-
True Knowledge Comes from Practice: Aligning LLMs with Embodied Environments via Reinforcement Learning 25 Jan 2024 · 1 repository · arXiv:2401.14151Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Linear Alignment: A Closed-form Solution for Aligning Human Preferences without Tuning and Feedback 21 Jan 2024 · 1 repository · arXiv:2401.11458Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
LangProp: A code optimization framework using Large Language Models applied to driving 18 Jan 2024 · 1 repository · arXiv:2401.10314Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
ReFT: Reasoning with Reinforced Fine-Tuning 17 Jan 2024 · 1 repository · arXiv:2401.08967Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 9 unverified (of 14 harvested samples) · 13 pointer-only (licence)
-
MAPO: Advancing Multilingual Reasoning through Multilingual Alignment-as-Preference Optimization 12 Jan 2024 · 1 repository · arXiv:2401.06838Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss 27 Dec 2023 · 1 repository · arXiv:2312.16682Syntology 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 8 harvested samples) · 1 pointer-only (licence)
-
Gradient Informed Proximal Policy Optimization 14 Dec 2023 · 1 repository · arXiv:2312.08710Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
DriveMLM: Aligning Multi-Modal Large Language Models with Behavioral Planning States for Autonomous Driving 14 Dec 2023 · 2 repositories · arXiv:2312.09245Syntology community repositories only · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Black-Box Prompt Optimization: Aligning Large Language Models without Model Training 7 Nov 2023 · 1 repository · arXiv:2311.04155Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples)
-
Reward Scale Robustness for Proximal Policy Optimization via DreamerV3 Tricks 26 Oct 2023 · 0 repositories · arXiv:2310.17805Syntology 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 1 pointer-only (licence)
-
Safe Navigation: Training Autonomous Vehicles using Deep Reinforcement Learning in CARLA 23 Oct 2023 · 1 repository · arXiv:2311.10735Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models 16 Oct 2023 · 3 repositories · arXiv:2310.10505Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 13 harvested samples) · 13 pointer-only (licence)
-
Distributional Soft Actor-Critic with Three Refinements 9 Oct 2023 · 2 repositories · arXiv:2310.05858Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Reinforcement Learning in the Era of LLMs: What is Essential? What is needed? An RL Perspective on RLHF, Prompting, and Beyond 9 Oct 2023 · 2 repositories · arXiv:2310.06147Syntology 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Reward Model Ensembles Help Mitigate Overoptimization 4 Oct 2023 · 2 repositories · arXiv:2310.02743Syntology official (archive's flag): 2 ran · 5 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 5 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 3 pointer-only (licence)
-
Cleanba: A Reproducible and Efficient Distributed Reinforcement Learning Platform 29 Sep 2023 · 1 repository · arXiv:2310.00036Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
Exploring the impact of low-rank adaptation on the performance, efficiency, and regularization of RLHF 16 Sep 2023 · 1 repository · arXiv:2309.09055Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 0 violated, 12 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 13 harvested samples)
-
Parallel Q-Learning: Scaling Off-policy Reinforcement Learning under Massively Parallel Simulation 24 Jul 2023 · 0 repositories · arXiv:2307.12983Syntology 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Llama 2: Open Foundation and Fine-Tuned Chat Models 18 Jul 2023 · 19 repositories · arXiv:2307.09288Syntology community repositories only · 33 ran (of which 7 constructed an object rather than computing a result; 22 with no instrument failure: 1 honoured, 1 violated, 20 with no contract checked; 11 where Syntology's instrument failed) · 19 unverified (of 52 harvested samples) · 20 pointer-only (licence)
-
Hidden Biases of End-to-End Driving Models 13 Jun 2023 · 1 repository · arXiv:2306.07957Syntology official (archive's flag): 13 ran · 13 ran (of which 10 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 3 where Syntology's instrument failed) · 4 unverified (of 17 harvested samples)
-
Fine-Tuning Language Models with Advantage-Induced Policy Alignment 4 Jun 2023 · 1 repository · arXiv:2306.02231Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples) · 1 pointer-only (licence)
-
Normalization Enhances Generalization in Visual Reinforcement Learning 1 Jun 2023 · 1 repository · arXiv:2306.00656Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Think Twice before Driving: Towards Scalable Decoders for End-to-End Autonomous Driving 10 May 2023 · 1 repository · arXiv:2305.06242Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples)
-
Reducing the Cost of Cycle-Time Tuning for Real-World Policy Optimization 9 May 2023 · 1 repository · arXiv:2305.05760Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 2 pointer-only (licence)
-
Local Optimization Achieves Global Optimality in Multi-Agent Reinforcement Learning 8 May 2023 · 1 repository · arXiv:2305.04819Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
RRHF: Rank Responses to Align Language Models with Human Feedback without tears 11 Apr 2023 · 1 repository · arXiv:2304.05302Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
AutoRL Hyperparameter Landscapes 5 Apr 2023 · 1 repository · arXiv:2304.02396Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples)
-
Behavior Proximal Policy Optimization 22 Feb 2023 · 2 repositories · arXiv:2302.11312Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Dynamic Simplex: Balancing Safety and Performance in Autonomous Cyber Physical Systems 20 Feb 2023 · 1 repository · arXiv:2302.09750Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Sample Dropout: A Simple yet Effective Variance Reduction Technique in Deep Policy Optimization 5 Feb 2023 · 1 repository · arXiv:2302.02299Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 6 unverified (of 12 harvested samples) · 12 pointer-only (licence)
-
Robust Policy Optimization in Deep Reinforcement Learning 14 Dec 2022 · 1 repository · arXiv:2212.07536Syntology official: no sample here; runs from other or unrecorded repositories · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Analyzing Infrastructure LiDAR Placement with Realistic LiDAR Simulation Library 29 Nov 2022 · 2 repositories · arXiv:2211.15975Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
PlanT: Explainable Planning Transformers via Object-Level Representations 25 Oct 2022 · 2 repositories · arXiv:2210.14222Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Model-Based Imitation Learning for Urban Driving 14 Oct 2022 · 1 repository · arXiv:2210.07729Syntology official (archive's flag): 21 ran · 21 ran (of which 13 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 0 violated, 13 with no contract checked; 8 where Syntology's instrument failed) · 2 unverified (of 23 harvested samples)
-
Real-Time Reinforcement Learning for Vision-Based Robotics Utilizing Local and Remote Computers 5 Oct 2022 · 2 repositories · arXiv:2210.02317Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
Learning to Generalize with Object-centric Agents in the Open World Survival Game Crafter 5 Aug 2022 · 1 repository · arXiv:2208.03374Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
Safety-Enhanced Autonomous Driving Using Interpretable Sensor Fusion Transformer 28 Jul 2022 · 1 repository · arXiv:2207.14024Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
ST-P3: End-to-end Vision-based Autonomous Driving via Spatial-Temporal Feature Learning 15 Jul 2022 · 1 repository · arXiv:2207.07601Syntology 5 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
Automated Detection of Label Errors in Semantic Segmentation Datasets via Deep Learning and Uncertainty Quantification 13 Jul 2022 · 1 repository · arXiv:2207.06104Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
EnvPool: A Highly Parallel Reinforcement Learning Environment Execution Engine 21 Jun 2022 · 3 repositories · arXiv:2206.10558Syntology official (archive's flag): 3 ran · 5 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 1 pointer-only (licence)
-
A Unified Approach to Reinforcement Learning, Quantal Response Equilibria, and Two-Player Zero-Sum Games 12 Jun 2022 · 3 repositories · arXiv:2206.05825Syntology official: no sample here; runs from other or unrecorded repositories · 5 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 1 honoured, 3 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
TransFuser: Imitation with Transformer-Based Sensor Fusion for Autonomous Driving 31 May 2022 · 3 repositories · arXiv:2205.15997Syntology community repositories only · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 1 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
Efficient Reward Poisoning Attacks on Online Deep Reinforcement Learning 30 May 2022 · 1 repository · arXiv:2205.14842Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Quark: Controllable Text Generation with Reinforced Unlearning 26 May 2022 · 1 repository · arXiv:2205.13636Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
An Evaluation Study of Intrinsic Motivation Techniques applied to Reinforcement Learning over Hard Exploration Environments 23 May 2022 · 1 repository · arXiv:2205.11184Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 2 pointer-only (licence)
-
Cliff Diving: Exploring Reward Surfaces in Reinforcement Learning Environments 14 May 2022 · 0 repositories · arXiv:2205.07015Syntology 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 4 harvested samples) · 1 pointer-only (licence)
-
KING: Generating Safety-Critical Driving Scenarios for Robust Imitation via Kinematics Gradients 28 Apr 2022 · 1 repository · arXiv:2204.13683Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 1 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples)
-
Learning from All Vehicles 22 Mar 2022 · 1 repository · arXiv:2203.11934Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples)
-
V2X-ViT: Vehicle-to-Everything Cooperative Perception with Vision Transformer 20 Mar 2022 · 1 repository · arXiv:2203.10638Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
PanoFlow: Learning 360° Optical Flow for Surrounding Temporal Understanding 27 Feb 2022 · 1 repository · arXiv:2202.13388Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
CADRE: A Cascade Deep Reinforcement Learning Framework for Vision-based Autonomous Urban Driving 17 Feb 2022 · 1 repository · arXiv:2202.08557Syntology official (archive's flag): 8 ran · 8 ran (of which 7 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 9 harvested samples)
-
Learning Distilled Collaboration Graph for Multi-Agent Perception 1 Nov 2021 · 2 repositories · arXiv:2111.00643Syntology community repositories only · 4 ran (of which 4 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified; every one of the 4 samples that ran constructed an object rather than computing a result (of 6 harvested samples)
-
Object-Aware Regularization for Addressing Causal Confusion in Imitation Learning 27 Oct 2021 · 1 repository · arXiv:2110.14118Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)