Methods › Reinforcement Learning › Offline Reinforcement Learning Methods › DPO › Papers where code ran, page 2
Direct Preference Optimization
DPO
Papers archive 2025-07-28
archive papers tagged: 409 · with a code link: 184 · where Syntology ran a sample: 112 (98 with a run with no instrument failure, 14 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (112 of 409 tagged: 98 with a run with no instrument failure, 14 where every run was a failure of Syntology's instrument)
Syntology We ran code from the paper's repository; we did not isolate this method inside it.
Page 2 of 2: papers 101 to 112 of the 112 tagged papers where Syntology ran at least one harvested sample (98 with a run with no instrument failure, 14 where every run was a failure of Syntology's instrument), newest first by the archive's date (ties by arXiv id). This is a filter on Syntology's record ordered by date only, not a ranking; a run is not a correctness claim. A paper missing from this list is not a recorded non-run: it may have no arXiv id, no harvested code, or only samples that have not run yet.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity 3 Jan 2024 · 2 repositories · arXiv:2401.01967Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 0 violated, 12 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 14 harvested samples)
-
Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss 27 Dec 2023 · 1 repository · arXiv:2312.16682Syntology 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 8 harvested samples) · 1 pointer-only (licence)
-
What Makes Good Data for Alignment? A Comprehensive Study of Automatic Data Selection in Instruction Tuning 25 Dec 2023 · 1 repository · arXiv:2312.15685Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples) · 1 pointer-only (licence)
-
Policy Optimization in RLHF: The Impact of Out-of-preference Data 17 Dec 2023 · 1 repository · arXiv:2312.10584Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
ULMA: Unified Language Model Alignment with Human Demonstration and Point-wise Preference 5 Dec 2023 · 1 repository · arXiv:2312.02554Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Diffusion Model Alignment Using Direct Preference Optimization 21 Nov 2023 · 2 repositories · arXiv:2311.12908Syntology 8 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 1 violated, 3 with no contract checked; 4 where Syntology's instrument failed) · 2 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
Direct Preference Optimization for Neural Machine Translation with Minimum Bayes Risk Decoding 14 Nov 2023 · 1 repository · arXiv:2311.08380Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 1 pointer-only (licence)
-
Black-Box Prompt Optimization: Aligning Large Language Models without Model Training 7 Nov 2023 · 1 repository · arXiv:2311.04155Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples)
-
Qilin-Med: Multi-stage Knowledge Injection Advanced Medical Large Language Model 13 Oct 2023 · 1 repository · arXiv:2310.09089Syntology official (archive's flag): 2 ran · 5 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Improving Summarization with Human Edits 9 Oct 2023 · 2 repositories · arXiv:2310.05857Syntology official (archive's flag): 3 ran · 3 ran (of which 1 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints 28 Sep 2023 · 1 repository · arXiv:2309.16240Syntology 16 ran (of which 0 constructed an object rather than computing a result; 15 with no instrument failure: 0 honoured, 0 violated, 15 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 17 harvested samples) · 1 pointer-only (licence)
-
Direct Preference Optimization: Your Language Model is Secretly a Reward Model 29 May 2023 · 29 repositories · arXiv:2305.18290Syntology 25 ran (of which 1 constructed an object rather than computing a result; 24 with no instrument failure: 0 honoured, 0 violated, 24 with no contract checked; 1 where Syntology's instrument failed) · 6 unverified (of 31 harvested samples) · 2 pointer-only (licence)