Preference Optimization by Estimating the Ratio of the Data Distribution
Yeongmin Kim, HeeSun Bae, Byeonghu Na, Il-Chul Moon
摘要
Direct preference optimization (DPO) is widely used as a simple and stable method for aligning large language models (LLMs) with human preferences. This paper investigates a generalized DPO loss that enables a policy model to match the target policy from a likelihood ratio estimation perspective. The ratio of the target policy provides a unique identification of the policy distribution without relying on reward models or partition functions. This allows the generalized loss to retain both simplicity and theoretical guarantees, which prior work such as -PO fails to achieve simultaneously. We propose Bregman preference optimization (BPO), a generalized framework for ratio matching that provides a family of objective functions achieving target policy optimality. BPO subsumes DPO as a special case and offers tractable forms for all instances, allowing implementation with a few lines of code. We further develop scaled Basu's power divergence (SBA), a gradient scaling method that can be used for BPO instances. The BPO framework complements other DPO variants and is applicable to target policies defined by these variants. In experiments, unlike other probabilistic loss extensions such as -DPO or -PO, which exhibit a trade-off between generation fidelity and diversity, instances of BPO improve both win rate and entropy compared with DPO. When applied to Llama-3-8B-Instruct, BPO achieves state-of-the-art performance among Llama-3-8B backbones, with a 55.9% length-controlled win rate on AlpacaEval2. Project page: https://github.com/aailab-kaist/BPO.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Lookahead Sample Reward Guidance for Test-Time Scaling of Diffusion ModelsYeongmin Kim, Donghyeok Shin, Byeonghu Na, Minsang Park 等ICML 2026 · 被引用 1 次
- TokenRatio: Principled Token-Level Preference Optimization via Ratio MatchingTruong Nguyen, Tien-Phat Nguyen, Linh Van, Duy Nguyen 等ICML 2026
- Towards Disentangled Preference Optimization Dynamics: Suppress the Loser, Preserve the WinnerWei Chen, Yubing Wu, Junmei Yang, Delu Zeng 等ICML 2026
它引用的顶会 Paper30
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
- SimPO: Simple Preference Optimization with a Reference-Free RewardYu Meng, Mengzhou Xia, Danqi ChenNeurIPS 2024 · 被引用 1,203 次
- Model Alignment as Prospect Theoretic OptimizationKawin Ethayarajh, Winnie Xu, Niklas Muennighoff, Dan Jurafsky 等ICML 2024 · 被引用 973 次
相关 Paper
- Normalized Rewards for Preference OptimizationShawn Im, Federico Danieli, Skyler Seto, Barry-John Theobald 等ICML 2026 · 被引用 571 次
- AlphaDPO: Adaptive Reward Margin for Direct Preference OptimizationJunkang Wu, Xue Wang, Zhengyi Yang, Jiancan Wu 等ICML 2025
- -SPPO: Semantic-Calibrated Self-Play Preference OptimizationXiwen Chen, Wenhui Zhu, Jingjing Wang, Peijie Qiu 等ICML 2026
- DiffPO: Diffusion-styled Preference Optimization for Inference Time Alignment of Large Language ModelsRuizhe Chen, Wenhao Chai, Zhifei Yang, Xiaotian Zhang 等ACL 2025
- Earlier Tokens Contribute More: Learning Direct Preference Optimization From Temporal Decay PerspectiveRuichen Shao, Bei Li, Gangao Liu, Yang Chen 等ICLR 2025
