Papers › Aligning LLMs with Domain Invariant Reward Models

Aligning LLMs with Domain Invariant Reward Models

1 Jan 2025arXiv:2501.00911archive 2025-07-28

David Wu, Sanjiban Choudhury

Aligning large language models (LLMs) to human preferences is challenging in domains where preference data is unavailable. We address the problem of learning reward models for such target domains by leveraging feedback collected from simpler source domains, where human preferences are easier to obtain. Our key insight is that, while domains may differ significantly, human preferences convey \emph{domain-agnostic} concepts that can be effectively captured by a reward model. We propose \method, a framework that trains domain-invariant reward models by optimizing a dual loss: a domain loss that minimizes the divergence between source and target distribution, and a source loss that optimizes preferences on the source domain. We show \method is a general approach that we evaluate and analyze across 4 distinct settings: (1) Cross-lingual transfer (accuracy: 0.621 →0.661), (2) Clean-to-noisy (accuracy: 0.671 →0.703), (3) Few-shot-to-full transfer (accuracy: 0.845 →0.920), and (4) Simple-to-complex tasks transfer (correlation: 0.508 →0.556). Our code, models and data are available at \url{https://github.com/portal-cornell/dial}.

PaperPDFCode

Code

portal-cornell/dial officialmentioned in papermentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Cross-Lingual Transfer

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections