Methods › Reinforcement Learning › Offline Reinforcement Learning Methods › DPO › Papers where code ran, page 1
Direct Preference Optimization
DPO
Papers archive 2025-07-28
archive papers tagged: 409 · with a code link: 184 · where Syntology ran a sample: 112 (98 with a run with no instrument failure, 14 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (112 of 409 tagged: 98 with a run with no instrument failure, 14 where every run was a failure of Syntology's instrument)
Syntology We ran code from the paper's repository; we did not isolate this method inside it.
Page 1 of 2: papers 1 to 100 of the 112 tagged papers where Syntology ran at least one harvested sample (98 with a run with no instrument failure, 14 where every run was a failure of Syntology's instrument), newest first by the archive's date (ties by arXiv id). This is a filter on Syntology's record ordered by date only, not a ranking; a run is not a correctness claim. A paper missing from this list is not a recorded non-run: it may have no arXiv id, no harvested code, or only samples that have not run yet.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
video-SALMONN 2: Captioning-Enhanced Audio-Visual Large Language Models 18 Jun 2025 · 1 repository · arXiv:2506.15220Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 8 harvested samples) · 1 pointer-only (licence)
-
TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization 17 Jun 2025 · 1 repository · arXiv:2506.14574Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
LeVo: High-Quality Song Generation with Multi-Preference Alignment 9 Jun 2025 · 1 repository · arXiv:2506.07520Syntology community repositories only · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
SuperWriter: Reflection-Driven Long-Form Generation with Large Language Models 4 Jun 2025 · 1 repository · arXiv:2506.04180Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Frictional Agent Alignment Framework: Slow Down and Don't Break Things 26 May 2025 · 1 repository · arXiv:2505.19428Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples)
-
Token-Importance Guided Direct Preference Optimization 26 May 2025 · 0 repositories · arXiv:2505.19653Syntology 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Understanding the Performance Gap in Preference Learning: A Dichotomy of RLHF and DPO 26 May 2025 · 1 repository · arXiv:2505.19770Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 4 harvested samples)
-
MPO: Multilingual Safety Alignment via Reward Gap Optimization 22 May 2025 · 1 repository · arXiv:2505.16869Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
MM-IFEngine: Towards Multimodal Instruction Following 10 Apr 2025 · 1 repository · arXiv:2504.07957Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 10 harvested samples)
-
ExCoT: Optimizing Reasoning for Text-to-SQL with Execution Feedback 25 Mar 2025 · 1 repository · arXiv:2503.19988Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 5 harvested samples)
-
TEMPLE:Temporal Preference Learning of Video LLMs via Difficulty Scheduling and Pre-SFT Alignment 21 Mar 2025 · 1 repository · arXiv:2503.16929Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond 13 Mar 2025 · 1 repository · arXiv:2503.10460Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 1 honoured, 1 violated, 6 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (of 13 harvested samples)
-
LightGen: Efficient Image Generation through Knowledge Distillation and Direct Preference Optimization 11 Mar 2025 · 1 repository · arXiv:2503.08619Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 3 honoured, 0 violated, 2 with no contract checked; 4 where Syntology's instrument failed) · 1 unverified (of 10 harvested samples) · 3 pointer-only (licence)
-
Improving LLM Safety Alignment with Dual-Objective Optimization 5 Mar 2025 · 1 repository · arXiv:2503.03710Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 1 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Can RLHF be More Efficient with Imperfect Reward Models? A Policy Coverage Perspective 26 Feb 2025 · 1 repository · arXiv:2502.19255Syntology community repositories only · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Agentic Reward Modeling: Integrating Human Preferences with Verifiable Correctness Signals for Reliable Reward Systems 26 Feb 2025 · 1 repository · arXiv:2502.19328Syntology official (archive's flag): 12 ran · 12 ran (of which 3 constructed an object rather than computing a result; 8 with no instrument failure: 4 honoured, 1 violated, 3 with no contract checked; 4 where Syntology's instrument failed) · 2 unverified (of 14 harvested samples)
-
HIPPO: Enhancing the Table Understanding Capability of Large Language Models through Hybrid-Modal Preference Optimization 24 Feb 2025 · 1 repository · arXiv:2502.17315Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Earlier Tokens Contribute More: Learning Direct Preference Optimization From Temporal Decay Perspective 20 Feb 2025 · 1 repository · arXiv:2502.14340Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 1 pointer-only (licence)
-
LongPO: Long Context Self-Evolution of Large Language Models through Short-to-Long Preference Optimization 19 Feb 2025 · 1 repository · arXiv:2502.13922Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model 14 Feb 2025 · 3 repositories · arXiv:2502.10248Syntology official (archive's flag): 1 ran · 9 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 9 harvested samples)
-
Principled Data Selection for Alignment: The Hidden Risks of Difficult Examples 11 Feb 2025 · 1 repository · arXiv:2502.09650Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
GuardReasoner: Towards Reasoning-based LLM Safeguards 30 Jan 2025 · 1 repository · arXiv:2501.18492Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
WILDCHAT-50M: A Deep Dive Into the Role of Synthetic Data in Post-Training 30 Jan 2025 · 1 repository · arXiv:2501.18511Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 9 harvested samples)
-
CHiP: Cross-modal Hierarchical Direct Preference Optimization for Multimodal LLMs 28 Jan 2025 · 1 repository · arXiv:2501.16629Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples)
-
Mitigating Hallucinations in Large Vision-Language Models via DPO: On-Policy Data Hold the Key 16 Jan 2025 · 1 repository · arXiv:2501.09695Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 2 where Syntology's instrument failed) · 5 unverified (of 16 harvested samples) · 16 pointer-only (licence)
-
Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding 14 Jan 2025 · 1 repository · arXiv:2501.07888Syntology official (archive's flag): 2 ran · 4 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
No Preference Left Behind: Group Distributional Preference Optimization 28 Dec 2024 · 1 repository · arXiv:2412.20299Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Offline Reinforcement Learning for LLM Multi-Step Reasoning 20 Dec 2024 · 2 repositories · arXiv:2412.16145Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Preference-Oriented Supervised Fine-Tuning: Favoring Target Model Over Aligned Large Language Models 17 Dec 2024 · 1 repository · arXiv:2412.12865Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
SPRec: Leveraging Self-Play to Debias Preference Alignment for Large Language Model-based Recommendations 12 Dec 2024 · 1 repository · arXiv:2412.09243Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
PatchDPO: Patch-level DPO for Finetuning-free Personalized Image Generation 4 Dec 2024 · 1 repository · arXiv:2412.03177Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 1 violated, 7 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning Capability 29 Nov 2024 · 1 repository · arXiv:2411.19943Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models 21 Nov 2024 · 1 repository · arXiv:2411.14432Syntology official (archive's flag): 3 ran · 5 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 5 where Syntology's instrument failed) · 5 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
Beyond Toxic Neurons: A Mechanistic Analysis of DPO for Toxicity Reduction 10 Nov 2024 · 1 repository · arXiv:2411.06424Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 0 violated, 13 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 15 harvested samples)
-
Sample-Efficient Alignment for LLMs 3 Nov 2024 · 1 repository · arXiv:2411.01493Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
TODO: Enhancing LLM Alignment with Ternary Preferences 2 Nov 2024 · 1 repository · arXiv:2411.02442Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 4 harvested samples)
-
VPO: Leveraging the Number of Votes in Preference Optimization 30 Oct 2024 · 1 repository · arXiv:2410.22891Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
f-PO: Generalizing Preference Optimization with f-divergence Minimization 29 Oct 2024 · 1 repository · arXiv:2410.21662Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 0 violated, 12 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 14 harvested samples)
-
LongReward: Improving Long-context Large Language Models with AI Feedback 28 Oct 2024 · 1 repository · arXiv:2410.21252Syntology official (archive's flag): 1 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 1 pointer-only (licence)
-
Fast Best-of-N Decoding via Speculative Rejection 26 Oct 2024 · 1 repository · arXiv:2410.20290Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
α-DPO: Adaptive Reward Margin is What Direct Preference Optimization Needs 14 Oct 2024 · 1 repository · arXiv:2410.10148Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 1 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 1 pointer-only (licence)
-
Language Imbalance Driven Rewarding for Multilingual Self-improving 11 Oct 2024 · 1 repository · arXiv:2410.08964Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 2 where Syntology's instrument failed) · 5 unverified (of 13 harvested samples) · 13 pointer-only (licence)
-
Reward-Augmented Data Enhances Direct Preference Alignment of LLMs 10 Oct 2024 · 1 repository · arXiv:2410.08067Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 6 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples)
-
TIS-DPO: Token-level Importance Sampling for Direct Preference Optimization With Estimated Weights 6 Oct 2024 · 2 repositories · arXiv:2410.04350Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
Regressing the Relative Future: Efficient Policy Optimization for Multi-turn RLHF 6 Oct 2024 · 1 repository · arXiv:2410.04612Syntology official (archive's flag): 14 ran · 14 ran (of which 0 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 0 violated, 13 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 15 harvested samples)
-
FlashMask: Efficient and Rich Mask Extension of FlashAttention 2 Oct 2024 · 1 repository · arXiv:2410.01359Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
The Crucial Role of Samplers in Online Direct Preference Optimization 29 Sep 2024 · 1 repository · arXiv:2409.19605Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
On Extending Direct Preference Optimization to Accommodate Ties 25 Sep 2024 · 0 repositories · arXiv:2409.17431Syntology 5 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 2 pointer-only (licence)
-
Instruct-SkillMix: A Powerful Pipeline for LLM Instruction Tuning 27 Aug 2024 · 1 repository · arXiv:2408.14774Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates 23 Aug 2024 · 1 repository · arXiv:2408.13006Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
Personality Alignment of Large Language Models 21 Aug 2024 · 1 repository · arXiv:2408.11779Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
xGen-MM (BLIP-3): A Family of Open Large Multimodal Models 16 Aug 2024 · 1 repository · arXiv:2408.08872Syntology 5 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Bridging and Modeling Correlations in Pairwise Data for Direct Preference Optimization 14 Aug 2024 · 1 repository · arXiv:2408.07471Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples) · 2 pointer-only (licence)
-
Self-Training with Direct Preference Optimization Improves Chain-of-Thought Reasoning 25 Jul 2024 · 1 repository · arXiv:2407.18248Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Understanding Reference Policies in Direct Preference Optimization 18 Jul 2024 · 1 repository · arXiv:2407.13709Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Learning Dynamics of LLM Finetuning 15 Jul 2024 · 1 repository · arXiv:2407.10490Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (of 10 harvested samples) · 2 pointer-only (licence)
-
Qwen2-Audio Technical Report 15 Jul 2024 · 2 repositories · arXiv:2407.10759Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
β-DPO: Direct Preference Optimization with Dynamic β 11 Jul 2024 · 1 repository · arXiv:2407.08639Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Towards Robust Alignment of Language Models: Distributionally Robustifying Direct Preference Optimization 10 Jul 2024 · 1 repository · arXiv:2407.07880Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Step-Controlled DPO: Leveraging Stepwise Error for Enhanced Mathematical Reasoning 30 Jun 2024 · 1 repository · arXiv:2407.00782Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
STLLaVA-Med: Self-Training Large Language and Vision Assistant for Medical Question-Answering 28 Jun 2024 · 1 repository · arXiv:2406.19973Syntology official (archive's flag): 9 ran · 9 ran (of which 4 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 4 where Syntology's instrument failed) · 10 unverified (of 19 harvested samples)
-
Decoding-Time Language Model Alignment with Multiple Objectives 27 Jun 2024 · 1 repository · arXiv:2406.18853Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample)
-
Suri: Multi-constraint Instruction Following for Long-form Text Generation 27 Jun 2024 · 1 repository · arXiv:2406.19371Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 1 violated, 11 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 15 harvested samples) · 15 pointer-only (licence)
-
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs 26 Jun 2024 · 1 repository · arXiv:2406.18629Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 12 harvested samples) · 12 pointer-only (licence)
-
Preference Tuning For Toxicity Mitigation Generalizes Across Languages 23 Jun 2024 · 1 repository · arXiv:2406.16235Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 9 harvested samples)
-
Direct Multi-Turn Preference Optimization for Language Agents 21 Jun 2024 · 1 repository · arXiv:2406.14868Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models 19 Jun 2024 · 1 repository · arXiv:2406.13542Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
DPO: Dual-Perturbation Optimization for Test-time Adaptation in 3D Object Detection 19 Jun 2024 · 1 repository · arXiv:2406.13891Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 5 harvested samples)
-
mDPO: Conditional Preference Optimization for Multimodal Large Language Models 17 Jun 2024 · 1 repository · arXiv:2406.11839Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 0 violated, 1 with no contract checked; 3 where Syntology's instrument failed) · 5 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
Is poisoning a real threat to LLM alignment? Maybe more so than you think 17 Jun 2024 · 1 repository · arXiv:2406.12091Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 1 pointer-only (licence)
-
Humor in AI: Massive Scale Crowd-Sourced Preferences and Benchmarks for Cartoon Captioning 15 Jun 2024 · 1 repository · arXiv:2406.10522Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 10 harvested samples)
-
On Softmax Direct Preference Optimization for Recommendation 13 Jun 2024 · 1 repository · arXiv:2406.09215Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback 13 Jun 2024 · 2 repositories · arXiv:2406.09279Syntology official (archive's flag): 15 ran · 15 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 7 where Syntology's instrument failed) · 4 unverified (of 19 harvested samples)
-
Triple Preference Optimization: Achieving Better Alignment with Less Data in a Single Step Optimization 26 May 2024 · 1 repository · arXiv:2405.16681Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples)
-
SimPO: Simple Preference Optimization with a Reference-Free Reward 23 May 2024 · 2 repositories · arXiv:2405.14734Syntology official (archive's flag): 5 ran · 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 1 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples) · 1 pointer-only (licence)
-
Curriculum Direct Preference Optimization for Diffusion and Consistency Models 22 May 2024 · 1 repository · arXiv:2405.13637Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 5 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
Model Editing as a Robust and Denoised variant of DPO: A Case Study on Toxicity 22 May 2024 · 2 repositories · arXiv:2405.13967Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 5 unverified (of 12 harvested samples)
-
Quantifying and Optimizing Global Faithfulness in Persona-driven Role-playing 13 May 2024 · 1 repository · arXiv:2405.07726Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 7 harvested samples)
-
Value Augmented Sampling for Language Model Alignment and Personalization 10 May 2024 · 1 repository · arXiv:2405.06639Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
D2PO: Discriminator-Guided DPO with Response Evaluation Models 2 May 2024 · 1 repository · arXiv:2405.01511Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 13 harvested samples) · 13 pointer-only (licence)
-
Self-Play Preference Optimization for Language Model Alignment 1 May 2024 · 1 repository · arXiv:2405.00675Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 1 pointer-only (licence)
-
DPO Meets PPO: Reinforced Token Optimization for RLHF 29 Apr 2024 · 1 repository · arXiv:2404.18922Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples) · 1 pointer-only (licence)
-
REBEL: Reinforcement Learning via Regressing Relative Rewards 25 Apr 2024 · 3 repositories · arXiv:2404.16767Syntology official (archive's flag): 16 ran · 16 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 1 honoured, 0 violated, 11 with no contract checked; 4 where Syntology's instrument failed) · 4 unverified (of 20 harvested samples) · 6 pointer-only (licence)
-
DPO: A Differential and Pointwise Control Approach to Reinforcement Learning 24 Apr 2024 · 0 repositories · arXiv:2404.15617Syntology 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample)
-
Filtered Direct Preference Optimization 22 Apr 2024 · 1 repository · arXiv:2404.13846Syntology official: harvested, nothing ran · 2 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified; every one of the 2 samples that ran constructed an object rather than computing a result (of 3 harvested samples)
-
Token-level Direct Preference Optimization 18 Apr 2024 · 1 repository · arXiv:2404.11999Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 1 pointer-only (licence)
-
Regularized Best-of-N Sampling with Minimum Bayes Risk Objective for Language Model Alignment 1 Apr 2024 · 1 repository · arXiv:2404.01054Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 10 harvested samples)
-
Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward 1 Apr 2024 · 1 repository · arXiv:2404.01258Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Comparing Bad Apples to Good Oranges: Aligning Large Language Models via Joint Preference Optimization 31 Mar 2024 · 1 repository · arXiv:2404.00530Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 11 harvested samples)
-
Disentangling Length from Quality in Direct Preference Optimization 28 Mar 2024 · 1 repository · arXiv:2403.19159Syntology 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Detoxifying Large Language Models via Knowledge Editing 21 Mar 2024 · 1 repository · arXiv:2403.14472Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 4 harvested samples)
-
Trial and Error: Exploration-Based Trajectory Optimization for LLM Agents 4 Mar 2024 · 2 repositories · arXiv:2403.02502Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 1 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
Direct Large Language Model Alignment Through Self-Rewarding Contrastive Prompt Distillation 19 Feb 2024 · 1 repository · arXiv:2402.11907Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Direct Preference Optimization with an Offset 16 Feb 2024 · 2 repositories · arXiv:2402.10571Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples) · 1 pointer-only (licence)
-
Multi-modal Preference Alignment Remedies Degradation of Visual Instruction Tuning on Language Models 16 Feb 2024 · 2 repositories · arXiv:2402.10884Syntology official (archive's flag): 4 ran · 5 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 1 honoured, 2 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Relative Preference Optimization: Enhancing LLM Alignment through Contrasting Responses across Identical and Diverse Prompts 12 Feb 2024 · 1 repository · arXiv:2402.10958Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Noise Contrastive Alignment of Language Models with Explicit Rewards 8 Feb 2024 · 3 repositories · arXiv:2402.05369Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
LiPO: Listwise Preference Optimization through Learning-to-Rank 2 Feb 2024 · 1 repository · arXiv:2402.01878Syntology 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples)
-
Towards Efficient Exact Optimization of Language Model Alignment 1 Feb 2024 · 2 repositories · arXiv:2402.00856Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
Self-Rewarding Language Models 18 Jan 2024 · 3 repositories · arXiv:2401.10020Syntology 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 2 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 5 unverified (of 11 harvested samples)