Lune

ICLR2026顶会

Displacement-Resistant Extensions of DPO with Nonconvex ff-Divergences

Idan Pipano, Shoham Sabach, Kavosh Asadi, Mohammad Ghavamzadeh

2026年份

摘要

DPO and related algorithms align language models by directly optimizing the RLHF objective: find a policy that maximizes the Bradley-Terry reward while staying close to a reference policy through a KL divergence penalty. Previous work showed that this approach could be further generalized: the original problem remains tractable even if the KL divergence is replaced by a family of ff-divergence with a convex generating function ff. Our first contribution is to show that convexity of ff is not essential. Instead, we identify a more general condition, referred to as DPO-inducing, that precisely characterizes when the RLHF problem remains tractable. Our next contribution is to establish a second condition on ff that is necessary to prevent probability displacement, a known empirical phenomenon in which the probabilities of the winner and the loser responses approach zero. We refer to any ff that satisfies this condition as displacement-resistant. We finally focus on a specific DPO-inducing and displacement-resistant ff, leading to our novel SquaredPO loss. Compared to DPO, this new loss offers stronger theoretical guarantees while performing competitively in practice.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext f44d271e-bd2f-451e-b4f0-6c2866eba3f4

它引用的顶会 Paper10

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖