Methods › Reinforcement Learning › Offline Reinforcement Learning Methods › DPO › Papers, page 2
Direct Preference Optimization
DPO
Papers archive 2025-07-28
archive papers tagged: 409 · with a code link: 184 · where Syntology ran a sample: 112 (98 with a run with no instrument failure, 14 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (112 of 409 tagged: 98 with a run with no instrument failure, 14 where every run was a failure of Syntology's instrument)
Page 2 of 5: papers 101 to 200 of 409, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Agentic Reward Modeling: Integrating Human Preferences with Verifiable Correctness Signals for Reliable Reward Systems 26 Feb 2025 · 1 repository · arXiv:2502.19328Syntology official (archive's flag): 12 ran · 12 ran (of which 3 constructed an object rather than computing a result; 8 with no instrument failure: 4 honoured, 1 violated, 3 with no contract checked; 4 where Syntology's instrument failed) · 2 unverified (of 14 harvested samples)
-
Can RLHF be More Efficient with Imperfect Reward Models? A Policy Coverage Perspective 26 Feb 2025 · 1 repository · arXiv:2502.19255Syntology community repositories only · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
OneRec: Unifying Retrieve and Rank with Generative Recommender and Iterative Preference Alignment 26 Feb 2025 · 0 repositories · arXiv:2502.18965
-
Debt Collection Negotiations with Large Language Models: An Evaluation System and Optimizing Decision Making with Multi-Agent 25 Feb 2025 · 0 repositories · arXiv:2502.18228
-
Aligning Compound AI Systems via System-level DPO 24 Feb 2025 · 0 repositories · arXiv:2502.17721
-
Finding the Sweet Spot: Preference Data Construction for Scaling Preference Optimization 24 Feb 2025 · 0 repositories · arXiv:2502.16825
-
HIPPO: Enhancing the Table Understanding Capability of Large Language Models through Hybrid-Modal Preference Optimization 24 Feb 2025 · 1 repository · arXiv:2502.17315Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
CODESYNC: Synchronizing Large Language Models with Dynamic Code Evolution at Scale 23 Feb 2025 · 1 repository · arXiv:2502.16645
-
C-3DPO: Constrained Controlled Classification for Direct Preference Optimization 22 Feb 2025 · 0 repositories · arXiv:2502.17507
-
Earlier Tokens Contribute More: Learning Direct Preference Optimization From Temporal Decay Perspective 20 Feb 2025 · 1 repository · arXiv:2502.14340Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 1 pointer-only (licence)
-
Federated Fine-Tuning of Large Language Models: Kahneman-Tversky vs. Direct Preference Optimization 20 Feb 2025 · 0 repositories · arXiv:2502.14187
-
Full-Step-DPO: Self-Supervised Preference Optimization with Step-wise Rewards for Mathematical Reasoning 20 Feb 2025 · 0 repositories · arXiv:2502.14356
-
Length-Controlled Margin-Based Preference Optimization without Reference Model 20 Feb 2025 · 1 repository · arXiv:2502.14643
-
Less is More: Improving LLM Alignment via Preference Data Selection 20 Feb 2025 · 0 repositories · arXiv:2502.14560
-
Self-Improvement Towards Pareto Optimality: Mitigating Preference Conflicts in Multi-Objective Alignment 20 Feb 2025 · 1 repository · arXiv:2502.14354
-
Efficient Safety Retrofitting Against Jailbreaking for LLMs 19 Feb 2025 · 0 repositories · arXiv:2502.13603
-
LongPO: Long Context Self-Evolution of Large Language Models through Short-to-Long Preference Optimization 19 Feb 2025 · 1 repository · arXiv:2502.13922Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
KL Penalty Control via Perturbation for Direct Preference Optimization 18 Feb 2025 · 1 repository · arXiv:2502.13177
-
Multi-Step Alignment as Markov Games: An Optimistic Online Gradient Descent Approach with Convergence Guarantees 18 Feb 2025 · 0 repositories · arXiv:2502.12678
-
CONSTRUCTA: Automating Commercial Construction Schedules in Fabrication Facilities with Large Language Models 17 Feb 2025 · 0 repositories · arXiv:2502.12066
-
Improve LLM-as-a-Judge Ability as a General Ability 17 Feb 2025 · 0 repositories · arXiv:2502.11689
-
Uncovering the Impact of Chain-of-Thought Reasoning for Direct Preference Optimization: Lessons from Text-to-SQL 17 Feb 2025 · 1 repository · arXiv:2502.11656
-
Balancing the Budget: Understanding Trade-offs Between Supervised and Preference-Based Finetuning 16 Feb 2025 · 0 repositories · arXiv:2502.11284
-
Preference learning made easy: Everything should be understood through win rate 14 Feb 2025 · 0 repositories · arXiv:2502.10505
-
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model 14 Feb 2025 · 3 repositories · arXiv:2502.10248Syntology official (archive's flag): 1 ran · 9 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 9 harvested samples)
-
DPO-Shift: Shifting the Distribution of Direct Preference Optimization 11 Feb 2025 · 1 repository · arXiv:2502.07599
-
Principled Data Selection for Alignment: The Hidden Risks of Difficult Examples 11 Feb 2025 · 1 repository · arXiv:2502.09650Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
Design Considerations in Offline Preference-based RL 8 Feb 2025 · 0 repositories · arXiv:2502.06861
-
LLMs Can Teach Themselves to Better Predict the Future 7 Feb 2025 · 1 repository · arXiv:2502.05253
-
Direct Distributional Optimization for Provable Alignment of Diffusion Models 5 Feb 2025 · 0 repositories · arXiv:2502.02954
-
DAMO: Data- and Model-aware Alignment of Multi-modal LLMs 4 Feb 2025 · 1 repository · arXiv:2502.01943
-
Distributionally Robust Direct Preference Optimization 4 Feb 2025 · 0 repositories · arXiv:2502.01930
-
Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models 4 Feb 2025 · 0 repositories · arXiv:2502.06812
-
LongDPO: Unlock Better Long-form Generation Abilities for LLMs via Critique-augmented Stepwise Information 4 Feb 2025 · 1 repository · arXiv:2502.02095
-
Eliciting Language Model Behaviors with Investigator Agents 3 Feb 2025 · 0 repositories · arXiv:2502.01236
-
The Differences Between Direct Alignment Algorithms are a Blur 3 Feb 2025 · 0 repositories · arXiv:2502.01237
-
HuViDPO:Enhancing Video Generation through Direct Preference Optimization for Human-Centric Alignment 2 Feb 2025 · 0 repositories · arXiv:2502.01690
-
Refining Alignment Framework for Diffusion Models with Intermediate-Step Preference Ranking 1 Feb 2025 · 0 repositories · arXiv:2502.01667
-
GuardReasoner: Towards Reasoning-based LLM Safeguards 30 Jan 2025 · 1 repository · arXiv:2501.18492Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
WILDCHAT-50M: A Deep Dive Into the Role of Synthetic Data in Post-Training 30 Jan 2025 · 1 repository · arXiv:2501.18511Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 9 harvested samples)
-
CHiP: Cross-modal Hierarchical Direct Preference Optimization for Multimodal LLMs 28 Jan 2025 · 1 repository · arXiv:2501.16629Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples)
-
Mitigating Hallucinations in Large Vision-Language Models via DPO: On-Policy Data Hold the Key 16 Jan 2025 · 1 repository · arXiv:2501.09695Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 2 where Syntology's instrument failed) · 5 unverified (of 16 harvested samples) · 16 pointer-only (licence)
-
Dynamic Portfolio Optimization via Augmented DDPG with Quantum Price Levels-Based Trading Strategy 15 Jan 2025 · 0 repositories · arXiv:2501.08528
-
Iterative Label Refinement Matters More than Preference Optimization under Weak Supervision 14 Jan 2025 · 1 repository · arXiv:2501.07886
-
Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding 14 Jan 2025 · 1 repository · arXiv:2501.07888Syntology official (archive's flag): 2 ran · 4 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
FocalPO: Enhancing Preference Optimizing by Focusing on Correct Preference Rankings 11 Jan 2025 · 0 repositories · arXiv:2501.06645
-
Personalized Preference Fine-tuning of Diffusion Models 11 Jan 2025 · 0 repositories · arXiv:2501.06655
-
Many of Your DPOs are Secretly One: Attempting Unification Through Mutual Information 2 Jan 2025 · 0 repositories · arXiv:2501.01544
-
RRHF-V: Ranking Responses to Mitigate Hallucinations in Multimodal Large Language Models with Human Feedback 1 Jan 2025 · 1 repository
-
Plug-and-Play Training Framework for Preference Optimization 30 Dec 2024 · 0 repositories · arXiv:2412.20996
-
No Preference Left Behind: Group Distributional Preference Optimization 28 Dec 2024 · 1 repository · arXiv:2412.20299Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Multimodal Preference Data Synthetic Alignment with Reward Model 23 Dec 2024 · 1 repository · arXiv:2412.17417
-
Understanding the Logic of Direct Preference Alignment through Logic 23 Dec 2024 · 0 repositories · arXiv:2412.17696
-
Teaching LLMs to Refine with Tools 22 Dec 2024 · 0 repositories · arXiv:2412.16871
-
Offline Reinforcement Learning for LLM Multi-Step Reasoning 20 Dec 2024 · 2 repositories · arXiv:2412.16145Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Northeastern Uni at Multilingual Counterspeech Generation: Enhancing Counter Speech Generation with LLM Alignment through Direct Preference Optimization 19 Dec 2024 · 0 repositories · arXiv:2412.15453
-
OnlineVPO: Align Video Diffusion Model with Online Video-Centric Preference Optimization 19 Dec 2024 · 0 repositories · arXiv:2412.15159
-
Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model 18 Dec 2024 · 0 repositories · arXiv:2412.13862
-
Preference-Oriented Supervised Fine-Tuning: Favoring Target Model Over Aligned Large Language Models 17 Dec 2024 · 1 repository · arXiv:2412.12865Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
SafetyDPO: Scalable Safety Alignment for Text-to-Image Generation 13 Dec 2024 · 0 repositories · arXiv:2412.10493
-
Sail into the Headwind: Alignment via Robust Rewards and Dynamic Labels against Reward Hacking 12 Dec 2024 · 0 repositories · arXiv:2412.09544
-
SPRec: Leveraging Self-Play to Debias Preference Alignment for Large Language Model-based Recommendations 12 Dec 2024 · 1 repository · arXiv:2412.09243Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
Knowledge Graph Guided Evaluation of Abstention Techniques 10 Dec 2024 · 0 repositories · arXiv:2412.07430
-
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation 8 Dec 2024 · 0 repositories · arXiv:2412.05818
-
Large Language Models for Ingredient Substitution in Food Recipes using Supervised Fine-tuning and Direct Preference Optimization 6 Dec 2024 · 1 repository · arXiv:2412.04922
-
SoPo: Text-to-Motion Generation Using Semi-Online Preference Optimization 6 Dec 2024 · 1 repository · arXiv:2412.05095
-
PatchDPO: Patch-level DPO for Finetuning-free Personalized Image Generation 4 Dec 2024 · 1 repository · arXiv:2412.03177Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 1 violated, 7 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning Capability 29 Nov 2024 · 1 repository · arXiv:2411.19943Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Mars-PO: Multi-Agent Reasoning System Preference Optimization 28 Nov 2024 · 0 repositories · arXiv:2411.19039
-
Improving the Transferability of Adversarial Attacks on Face Recognition with Diverse Parameters Augmentation 23 Nov 2024 · 0 repositories · arXiv:2411.15555
-
Reward Fine-Tuning Two-Step Diffusion Models via Learning Differentiable Latent-Space Surrogate Reward 22 Nov 2024 · 0 repositories · arXiv:2411.15247
-
Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models 21 Nov 2024 · 1 repository · arXiv:2411.14432Syntology official (archive's flag): 3 ran · 5 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 5 where Syntology's instrument failed) · 5 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
Entropy Controllable Direct Preference Optimization 12 Nov 2024 · 0 repositories · arXiv:2411.07595
-
Beyond Toxic Neurons: A Mechanistic Analysis of DPO for Toxicity Reduction 10 Nov 2024 · 1 repository · arXiv:2411.06424Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 0 violated, 13 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 15 harvested samples)
-
Learning Loss Landscapes in Preference Optimization 10 Nov 2024 · 0 repositories · arXiv:2411.06568
-
IOPO: Empowering LLMs with Complex Instruction Following via Input-Output Preference Optimization 9 Nov 2024 · 1 repository · arXiv:2411.06208
-
Abstract2Appendix: Academic Reviews Enhance LLM Long-Context Capabilities 7 Nov 2024 · 1 repository · arXiv:2411.05232
-
Towards Improved Preference Optimization Pipeline: from Data Generation to Budget-Controlled Regularization 7 Nov 2024 · 0 repositories · arXiv:2411.05875
-
SEE-DPO: Self Entropy Enhanced Direct Preference Optimization 6 Nov 2024 · 0 repositories · arXiv:2411.04712
-
The Root Shapes the Fruit: On the Persistence of Gender-Exclusive Harms in Aligned Language Models 6 Nov 2024 · 0 repositories · arXiv:2411.03700
-
Sample-Efficient Alignment for LLMs 3 Nov 2024 · 1 repository · arXiv:2411.01493Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
TODO: Enhancing LLM Alignment with Ternary Preferences 2 Nov 2024 · 1 repository · arXiv:2411.02442Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 4 harvested samples)
-
Scalable Reinforcement Post-Training Beyond Static Human Prompts: Evolving Alignment via Asymmetric Self-Play 31 Oct 2024 · 0 repositories · arXiv:2411.00062
-
VPO: Leveraging the Number of Votes in Preference Optimization 30 Oct 2024 · 1 repository · arXiv:2410.22891Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
f-PO: Generalizing Preference Optimization with f-divergence Minimization 29 Oct 2024 · 1 repository · arXiv:2410.21662Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 0 violated, 12 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 14 harvested samples)
-
Flow-DPO: Improving LLM Mathematical Reasoning through Online Multi-Agent Learning 29 Oct 2024 · 0 repositories · arXiv:2410.22304
-
LongReward: Improving Long-context Large Language Models with AI Feedback 28 Oct 2024 · 1 repository · arXiv:2410.21252Syntology official (archive's flag): 1 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 1 pointer-only (licence)
-
Accelerating Direct Preference Optimization with Prefix Sharing 27 Oct 2024 · 1 repository · arXiv:2410.20305
-
Learning from Response not Preference: A Stackelberg Approach for LLM Detoxification using Non-parallel Data 27 Oct 2024 · 1 repository · arXiv:2410.20298
-
Fast Best-of-N Decoding via Speculative Rejection 26 Oct 2024 · 1 repository · arXiv:2410.20290Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Uncertainty-Penalized Direct Preference Optimization 26 Oct 2024 · 0 repositories · arXiv:2410.20187
-
2D-DPO: Scaling Direct Preference Optimization with 2-Dimensional Supervision 25 Oct 2024 · 0 repositories · arXiv:2410.19720
-
Improving Inverse Folding for Peptide Design with Diversity-regularized Direct Preference Optimization 25 Oct 2024 · 0 repositories · arXiv:2410.19471
-
Aligning CodeLLMs with Direct Preference Optimization 24 Oct 2024 · 0 repositories · arXiv:2410.18585
-
Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models 23 Oct 2024 · 1 repository · arXiv:2410.18252Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
Scalable Ranked Preference Optimization for Text-to-Image Generation 23 Oct 2024 · 0 repositories · arXiv:2410.18013
-
AdvAgent: Controllable Blackbox Red-teaming on Web Agents 22 Oct 2024 · 0 repositories · arXiv:2410.17401
-
Optimizing LLMs with Direct Preferences: A Data Efficiency Perspective 22 Oct 2024 · 0 repositories · arXiv:2410.16586
-
A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications 21 Oct 2024 · 0 repositories · arXiv:2410.15595
-
GDPO: Learning to Directly Align Language Models with Diversity Using GFlowNets 19 Oct 2024 · 0 repositories · arXiv:2410.15096