Papers › RePO: ReLU-based Preference Optimization

RePO: ReLU-based Preference Optimization

10 Mar 2025arXiv:2503.07426archive 2025-07-28

Junkang Wu, Kexin Huang, Xue Wang, Jinyang Gao, Bolin Ding, Jiancan Wu, Xiangnan He, Xiang Wang

Aligning large language models (LLMs) with human preferences is critical for real-world deployment, yet existing methods like RLHF face computational and stability challenges. While DPO establishes an offline paradigm with single hyperparameter β, subsequent methods like SimPO reintroduce complexity through dual parameters (β, γ). We propose {ReLU-based Preference Optimization (RePO)}, a streamlined algorithm that eliminates β via two advances: (1) retaining SimPO's reference-free margins but removing β through gradient analysis, and (2) adopting a ReLU-based max-margin loss that naturally filters trivial pairs. Theoretically, RePO is characterized as SimPO's limiting case (β→∞), where the logistic weighting collapses to binary thresholding, forming a convex envelope of the 0-1 loss. Empirical results on AlpacaEval 2 and Arena-Hard show that RePO outperforms DPO and SimPO across multiple base models, requiring only one hyperparameter to tune.

PaperPDFCode

Code

junkangwu/repo officialmentioned in paperpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Methods

BASEDPO

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections