Methods › Reinforcement Learning › Offline Reinforcement Learning Methods › DPO › Papers, page 3
Direct Preference Optimization
DPO
Papers archive 2025-07-28
archive papers tagged: 409 · with a code link: 184 · where Syntology ran a sample: 112 (98 with a run with no instrument failure, 14 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (112 of 409 tagged: 98 with a run with no instrument failure, 14 where every run was a failure of Syntology's instrument)
Page 3 of 5: papers 201 to 300 of 409, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
DeformPAM: Data-Efficient Learning for Long-horizon Deformable Object Manipulation via Preference-based Action Alignment 15 Oct 2024 · 1 repository · arXiv:2410.11584
-
Have the VLMs Lost Confidence? A Study of Sycophancy in VLMs 15 Oct 2024 · 0 repositories · arXiv:2410.11302
-
α-DPO: Adaptive Reward Margin is What Direct Preference Optimization Needs 14 Oct 2024 · 1 repository · arXiv:2410.10148Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 1 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 1 pointer-only (licence)
-
Enhancing Multi-Step Reasoning Abilities of Language Models through Direct Q-Function Optimization 11 Oct 2024 · 2 repositories · arXiv:2410.09302
-
Language Imbalance Driven Rewarding for Multilingual Self-improving 11 Oct 2024 · 1 repository · arXiv:2410.08964Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 2 where Syntology's instrument failed) · 5 unverified (of 13 harvested samples) · 13 pointer-only (licence)
-
Simultaneous Reward Distillation and Preference Learning: Get You a Language Model Who Can Do Both 11 Oct 2024 · 0 repositories · arXiv:2410.08458
-
SuperCorrect: Supervising and Correcting Language Models with Error-Driven Insights 11 Oct 2024 · 2 repositories · arXiv:2410.09008Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Evolutionary Contrastive Distillation for Language Model Alignment 10 Oct 2024 · 0 repositories · arXiv:2410.07513
-
COS-DPO: Conditioned One-Shot Multi-Objective Fine-Tuning Framework 10 Oct 2024 · 0 repositories · arXiv:2410.08316
-
Optima: Optimizing Effectiveness and Efficiency for LLM-Based Multi-Agent System 10 Oct 2024 · 0 repositories · arXiv:2410.08115
-
Reward-Augmented Data Enhances Direct Preference Alignment of LLMs 10 Oct 2024 · 1 repository · arXiv:2410.08067Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 6 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples)
-
TPO: Aligning Large Language Models with Multi-branch & Multi-step Preference Trees 10 Oct 2024 · 0 repositories · arXiv:2410.12854
-
Enhancing Multimodal LLM for Detailed and Accurate Video Captioning using Multi-Round Preference Optimization 9 Oct 2024 · 0 repositories · arXiv:2410.06682
-
Subtle Errors Matter: Preference Learning via Error-injected Self-editing 9 Oct 2024 · 0 repositories · arXiv:2410.06638
-
Towards Self-Improvement of LLMs via MCTS: Leveraging Stepwise Knowledge with Curriculum Preference Learning 9 Oct 2024 · 0 repositories · arXiv:2410.06508
-
Accelerated Preference Optimization for Large Language Model Alignment 8 Oct 2024 · 0 repositories · arXiv:2410.06293
-
Direct Preference Optimization for LLM-Enhanced Recommendation Systems 8 Oct 2024 · 0 repositories · arXiv:2410.05939
-
As Simple as Fine-tuning: LLM Alignment via Bidirectional Negative Feedback Loss 7 Oct 2024 · 0 repositories · arXiv:2410.04834
-
Self-rationalization improves LLM as a fine-grained judge 7 Oct 2024 · 0 repositories · arXiv:2410.05495
-
Regressing the Relative Future: Efficient Policy Optimization for Multi-turn RLHF 6 Oct 2024 · 1 repository · arXiv:2410.04612Syntology official (archive's flag): 14 ran · 14 ran (of which 0 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 0 violated, 13 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 15 harvested samples)
-
TIS-DPO: Token-level Importance Sampling for Direct Preference Optimization With Estimated Weights 6 Oct 2024 · 2 repositories · arXiv:2410.04350Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
RainbowPO: A Unified Framework for Combining Improvements in Preference Optimization 5 Oct 2024 · 0 repositories · arXiv:2410.04203
-
An Exploration of Self-Supervised Mutual Information Alignment for Multi-Task Settings 2 Oct 2024 · 1 repository · arXiv:2410.01704
-
FlashMask: Efficient and Rich Mask Extension of FlashAttention 2 Oct 2024 · 1 repository · arXiv:2410.01359Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
The Perfect Blend: Redefining RLHF with Mixture of Judges 30 Sep 2024 · 0 repositories · arXiv:2409.20370
-
The Crucial Role of Samplers in Online Direct Preference Optimization 29 Sep 2024 · 1 repository · arXiv:2409.19605Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
Cross-lingual Human-Preference Alignment for Neural Machine Translation with Direct Quality Optimization 26 Sep 2024 · 0 repositories · arXiv:2409.17673
-
Modulated Intervention Preference Optimization (MIPO): Keep the Easy, Refine the Difficult 26 Sep 2024 · 0 repositories · arXiv:2409.17545
-
On Extending Direct Preference Optimization to Accommodate Ties 25 Sep 2024 · 0 repositories · arXiv:2409.17431Syntology 5 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 2 pointer-only (licence)
-
Zeroth-Order Policy Gradient for Reinforcement Learning from Human Feedback without Reward Inference 25 Sep 2024 · 0 repositories · arXiv:2409.17401
-
Orthogonal Finetuning for Direct Preference Optimization 23 Sep 2024 · 0 repositories · arXiv:2409.14836
-
Backtracking Improves Generation Safety 22 Sep 2024 · 0 repositories · arXiv:2409.14586
-
RRM: Robust Reward Model Training Mitigates Reward Hacking 20 Sep 2024 · 0 repositories · arXiv:2409.13156
-
Fine Tuning Large Language Models for Medicine: The Role and Importance of Direct Preference Optimization 19 Sep 2024 · 0 repositories · arXiv:2409.12741
-
From Lists to Emojis: How Format Bias Affects Model Alignment 18 Sep 2024 · 0 repositories · arXiv:2409.11704
-
ASFT: Aligned Supervised Fine-Tuning through Absolute Likelihood 14 Sep 2024 · 1 repository · arXiv:2409.10571
-
Geometric-Averaged Preference Optimization for Soft Preference Labels 10 Sep 2024 · 0 repositories · arXiv:2409.06691
-
Length Desensitization in Direct Preference Optimization 10 Sep 2024 · 0 repositories · arXiv:2409.06411
-
On the Limited Generalization Capability of the Implicit Reward Model Induced by Direct Preference Optimization 5 Sep 2024 · 0 repositories · arXiv:2409.03650
-
Building Math Agents with Multi-Turn Iterative Preference Learning 4 Sep 2024 · 0 repositories · arXiv:2409.02392
-
A Gradient Analysis Framework for Rewarding Good and Penalizing Bad Examples in Language Models 29 Aug 2024 · 0 repositories · arXiv:2408.16751
-
An Extremely Data-efficient and Generative LLM-based Reinforcement Learning Agent for Recommenders 28 Aug 2024 · 0 repositories · arXiv:2408.16032
-
Generative Verifiers: Reward Modeling as Next-Token Prediction 27 Aug 2024 · 0 repositories · arXiv:2408.15240
-
Instruct-SkillMix: A Powerful Pipeline for LLM Instruction Tuning 27 Aug 2024 · 1 repository · arXiv:2408.14774Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
UNA: Unifying Alignments of RLHF/PPO, DPO and KTO by a Generalized Implicit Reward Function 27 Aug 2024 · 0 repositories · arXiv:2408.15339
-
Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates 23 Aug 2024 · 1 repository · arXiv:2408.13006Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
Personality Alignment of Large Language Models 21 Aug 2024 · 1 repository · arXiv:2408.11779Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
Minor SFT loss for LLM fine-tune to increase performance and reduce model deviation 20 Aug 2024 · 0 repositories · arXiv:2408.10642
-
Minor DPO reject penalty to increase training robustness 19 Aug 2024 · 0 repositories · arXiv:2408.09834
-
xGen-MM (BLIP-3): A Family of Open Large Multimodal Models 16 Aug 2024 · 1 repository · arXiv:2408.08872Syntology 5 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Bridging and Modeling Correlations in Pairwise Data for Direct Preference Optimization 14 Aug 2024 · 1 repository · arXiv:2408.07471Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples) · 2 pointer-only (licence)
-
LongWriter: Unleashing 10,000+ Word Generation from Long Context LLMs 13 Aug 2024 · 3 repositories · arXiv:2408.07055
-
FuxiTranyu: A Multilingual Large Language Model Trained with Balanced Data 12 Aug 2024 · 1 repository · arXiv:2408.06273
-
A Logical Fallacy-Informed Framework for Argument Generation 7 Aug 2024 · 1 repository · arXiv:2408.03618
-
Intermediate direct preference optimization 6 Aug 2024 · 0 repositories · arXiv:2408.02923
-
On the Generalization of Preference Learning with DPO 6 Aug 2024 · 0 repositories · arXiv:2408.03459
-
Self-Training with Direct Preference Optimization Improves Chain-of-Thought Reasoning 25 Jul 2024 · 1 repository · arXiv:2407.18248Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
The Hitchhiker's Guide to Human Alignment with *PO 21 Jul 2024 · 0 repositories · arXiv:2407.15229
-
Clinical Reading Comprehension with Encoder-Decoder Models Enhanced by Direct Preference Optimization 19 Jul 2024 · 0 repositories · arXiv:2407.14000
-
Decomposed Direct Preference Optimization for Structure-Based Drug Design 19 Jul 2024 · 0 repositories · arXiv:2407.13981
-
Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization 18 Jul 2024 · 0 repositories · arXiv:2407.13399
-
Understanding Reference Policies in Direct Preference Optimization 18 Jul 2024 · 1 repository · arXiv:2407.13709Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
PersLLM: A Personified Training Approach for Large Language Models 17 Jul 2024 · 1 repository · arXiv:2407.12393
-
Learning Dynamics of LLM Finetuning 15 Jul 2024 · 1 repository · arXiv:2407.10490Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (of 10 harvested samples) · 2 pointer-only (licence)
-
Qwen2-Audio Technical Report 15 Jul 2024 · 2 repositories · arXiv:2407.10759Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
New Desiderata for Direct Preference Optimization 12 Jul 2024 · 0 repositories · arXiv:2407.09072
-
β-DPO: Direct Preference Optimization with Dynamic β 11 Jul 2024 · 1 repository · arXiv:2407.08639Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Towards Robust Alignment of Language Models: Distributionally Robustifying Direct Preference Optimization 10 Jul 2024 · 1 repository · arXiv:2407.07880Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Efficient and Accurate Memorable Conversation Model using DPO based on sLLM 9 Jul 2024 · 0 repositories · arXiv:2407.06537
-
LIONs: An Empirically Optimized Approach to Align Language Models 9 Jul 2024 · 1 repository · arXiv:2407.06542
-
Exposing Privacy Gaps: Membership Inference Attack on Preference Data for LLM Alignment 8 Jul 2024 · 0 repositories · arXiv:2407.06443
-
Towards Human Understanding of Paraphrase Types in ChatGPT 2 Jul 2024 · 3 repositories · arXiv:2407.02302
-
Learning to Explore and Select for Coverage-Conditioned Retrieval-Augmented Generation 1 Jul 2024 · 1 repository · arXiv:2407.01158
-
Explaining Length Bias in LLM-Based Preference Evaluations 1 Jul 2024 · 0 repositories · arXiv:2407.01085
-
Step-Controlled DPO: Leveraging Stepwise Error for Enhanced Mathematical Reasoning 30 Jun 2024 · 1 repository · arXiv:2407.00782Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
STLLaVA-Med: Self-Training Large Language and Vision Assistant for Medical Question-Answering 28 Jun 2024 · 1 repository · arXiv:2406.19973Syntology official (archive's flag): 9 ran · 9 ran (of which 4 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 4 where Syntology's instrument failed) · 10 unverified (of 19 harvested samples)
-
Decoding-Time Language Model Alignment with Multiple Objectives 27 Jun 2024 · 1 repository · arXiv:2406.18853Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample)
-
Suri: Multi-constraint Instruction Following for Long-form Text Generation 27 Jun 2024 · 1 repository · arXiv:2406.19371Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 1 violated, 11 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 15 harvested samples) · 15 pointer-only (licence)
-
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs 26 Jun 2024 · 1 repository · arXiv:2406.18629Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 12 harvested samples) · 12 pointer-only (licence)
-
Learning to Ask Informative Questions: Enhancing LLMs with Preference Optimization and Expected Information Gain 25 Jun 2024 · 1 repository · arXiv:2406.17453Syntology official: harvested, nothing ran · 0 ran · 5 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Not All Preference Pairs Are Created Equal: A Recipe for Annotation-Efficient Iterative Preference Learning 25 Jun 2024 · 0 repositories · arXiv:2406.17312
-
PAFT: A Parallel Training Paradigm for Effective LLM Fine-Tuning 25 Jun 2024 · 0 repositories · arXiv:2406.17923
-
Preference Tuning For Toxicity Mitigation Generalizes Across Languages 23 Jun 2024 · 1 repository · arXiv:2406.16235Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 9 harvested samples)
-
Direct Multi-Turn Preference Optimization for Language Agents 21 Jun 2024 · 1 repository · arXiv:2406.14868Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
SAIL: Self-Improving Efficient Online Alignment of Large Language Models 21 Jun 2024 · 0 repositories · arXiv:2406.15567
-
DPO: Dual-Perturbation Optimization for Test-time Adaptation in 3D Object Detection 19 Jun 2024 · 1 repository · arXiv:2406.13891Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 5 harvested samples)
-
Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models 19 Jun 2024 · 1 repository · arXiv:2406.13542Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Is poisoning a real threat to LLM alignment? Maybe more so than you think 17 Jun 2024 · 1 repository · arXiv:2406.12091Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 1 pointer-only (licence)
-
Iterative Length-Regularized Direct Preference Optimization: A Case Study on Improving 7B Language Models to GPT-4 Level 17 Jun 2024 · 0 repositories · arXiv:2406.11817
-
mDPO: Conditional Preference Optimization for Multimodal Large Language Models 17 Jun 2024 · 1 repository · arXiv:2406.11839Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 0 violated, 1 with no contract checked; 3 where Syntology's instrument failed) · 5 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
Eliminating Biased Length Reliance of Direct Preference Optimization via Down-Sampled KL Divergence 16 Jun 2024 · 1 repository · arXiv:2406.10957
-
Humor in AI: Massive Scale Crowd-Sourced Preferences and Benchmarks for Cartoon Captioning 15 Jun 2024 · 1 repository · arXiv:2406.10522Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 10 harvested samples)
-
Bootstrapping Language Models with DPO Implicit Rewards 14 Jun 2024 · 1 repository · arXiv:2406.09760Syntology official: harvested, nothing ran · 0 ran · 2 unverified (of 2 harvested samples)
-
Knowledge Editing in Language Models via Adapted Direct Preference Optimization 14 Jun 2024 · 0 repositories · arXiv:2406.09920
-
On Softmax Direct Preference Optimization for Recommendation 13 Jun 2024 · 1 repository · arXiv:2406.09215Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback 13 Jun 2024 · 2 repositories · arXiv:2406.09279Syntology official (archive's flag): 15 ran · 15 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 7 where Syntology's instrument failed) · 4 unverified (of 19 harvested samples)
-
3D-Properties: Identifying Challenges in DPO and Charting a Path Forward 11 Jun 2024 · 0 repositories · arXiv:2406.07327
-
Learning Reward and Policy Jointly from Demonstration and Preference Improves Alignment 11 Jun 2024 · 0 repositories · arXiv:2406.06874
-
Direct Preference Optimization for Suppressing Hallucinated Prior Exams in Radiology Report Generation 10 Jun 2024 · 0 repositories · arXiv:2406.06496
-
Margin-aware Preference Optimization for Aligning Diffusion Models without Reference 10 Jun 2024 · 0 repositories · arXiv:2406.06424