Methods › Reinforcement Learning › Offline Reinforcement Learning Methods › DPO
Direct Preference Optimization
DPO
Introduced by Rafael Rafailov et al. in Direct Preference Optimization: Your Language Model is Secretly a Reward Model
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
The archive carries no description for this method.
Papers archive 2025-07-28
30 shown of 409, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
Multi-Preference Lambda-weighted Listwise DPO for Dynamic Preference Alignment 24 Jun 2025 · 1 repository · arXiv:2506.19780
-
Smart-LLaMA-DPO: Reinforced Large Language Model for Explainable Smart Contract Vulnerability Detection 23 Jun 2025 · 0 repositories · arXiv:2506.18245
-
video-SALMONN 2: Captioning-Enhanced Audio-Visual Large Language Models 18 Jun 2025 · 1 repository · arXiv:2506.15220Syntology ran 3 of 8 samples · 5 unverified
-
TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization 17 Jun 2025 · 1 repository · arXiv:2506.14574Syntology ran 2 of 3 samples · 1 unverified · 3 pointer-only (licence)
-
Alignment Quality Index (AQI) : Beyond Refusals: AQI as an Intrinsic Alignment Diagnostic via Latent Geometry, Cluster Divergence, and Layer wise Pooled Representations 16 Jun 2025 · 0 repositories · arXiv:2506.13901
-
CAMS: A CityGPT-Powered Agentic Framework for Urban Human Mobility Simulation 16 Jun 2025 · 0 repositories · arXiv:2506.13599
-
From Judgment to Interference: Early Stopping LLM Harmful Outputs via Streaming Content Monitoring 11 Jun 2025 · 0 repositories · arXiv:2506.09996
-
Towards Bridging the Reward-Generation Gap in Direct Alignment Algorithms 11 Jun 2025 · 0 repositories · arXiv:2506.09457
-
Vision Matters: Simple Visual Perturbations Can Boost Multimodal Math Reasoning 11 Jun 2025 · 1 repository · arXiv:2506.09736
-
Explicit Preference Optimization: No Need for an Implicit Reward Model 9 Jun 2025 · 1 repository · arXiv:2506.07492
-
LeVo: High-Quality Song Generation with Multi-Preference Alignment 9 Jun 2025 · 1 repository · arXiv:2506.07520Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)
-
QA-LIGN: Aligning LLMs through Constitutionally Decomposed QA 9 Jun 2025 · 0 repositories · arXiv:2506.08123
-
LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs 5 Jun 2025 · 1 repository · arXiv:2506.05260
-
Aligning Large Language Models with Implicit Preferences from User-Generated Content 4 Jun 2025 · 0 repositories · arXiv:2506.04463
-
SuperWriter: Reflection-Driven Long-Form Generation with Large Language Models 4 Jun 2025 · 1 repository · arXiv:2506.04180Syntology ran 6 of 6 samples · 0 unverified · 6 pointer-only (licence)
-
Protein Inverse Folding From Structure Feedback 3 Jun 2025 · 0 repositories · arXiv:2506.03028
-
Smoothed Preference Optimization via ReNoise Inversion for Aligning Diffusion Models with Varied Human Preferences 3 Jun 2025 · 0 repositories · arXiv:2506.02698
-
Understanding the Impact of Sampling Quality in Direct Preference Optimization 3 Jun 2025 · 0 repositories · arXiv:2506.04272
-
IF-GUIDE: Influence Function-Guided Detoxification of LLMs 2 Jun 2025 · 1 repository · arXiv:2506.01790Syntology ran 0 of 1 samples · 1 unverified
-
Generalizable LLM Learning of Graph Synthetic Data with Reinforcement Learning 1 Jun 2025 · 0 repositories · arXiv:2506.00845
-
Harnessing Negative Signals: Reinforcement Distillation from Teacher Data for LLM Reasoning 30 May 2025 · 1 repository · arXiv:2505.24850
-
Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization 29 May 2025 · 0 repositories · arXiv:2505.23387
-
Differential Information: An Information-Theoretic Perspective on Preference Optimization 29 May 2025 · 0 repositories · arXiv:2505.23761
-
Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? 29 May 2025 · 0 repositories · arXiv:2505.23749
-
Improving Multilingual Social Media Insights: Aspect-based Comment Analysis 29 May 2025 · 0 repositories · arXiv:2505.23037
-
MCP Safety Training: Learning to Refuse Falsely Benign MCP Exploits using Improved Preference Alignment 29 May 2025 · 0 repositories · arXiv:2505.23634
-
Proximalized Preference Optimization for Diverse Feedback Types: A Decomposed Perspective on DPO 29 May 2025 · 0 repositories · arXiv:2505.23316
-
Reinforcement Learning for Better Verbalized Confidence in Long-Form Generation 29 May 2025 · 0 repositories · arXiv:2505.23912
-
D-Fusion: Direct Preference Optimization for Aligning Diffusion Models with Visually Consistent Samples 28 May 2025 · 0 repositories · arXiv:2505.22002
-
Enhancing Paraphrase Type Generation: The Impact of DPO and RLHF Evaluated with Human-Ranked Data 28 May 2025 · 1 repository · arXiv:2506.02018
Tasks archive 2025-07-28
20 shown of 181 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections