On Extending Direct Preference Optimization to Accommodate Ties
Jinghong Chen, Guangyu Yang, Weizhe Lin, Jingbiao Mei, Chenxu Lyu, Bill Byrne
摘要
We derive and investigate two DPO variants that explicitly model the possibility of declaring a tie in pair-wise comparisons. We replace the Bradley-Terry model in DPO with two well-known modeling extensions, by Rao and Kupper and by Davidson, that assign probability to ties as alternatives to clear preferences. Our experiments in neural machine translation and summarization show that explicitly labeled ties can be added to the datasets for these DPO variants without the degradation in task performance that is observed when the same tied pairs are presented to DPO. We find empirically that the inclusion of ties leads to stronger regularization with respect to the reference policy as measured by KL divergence, and we see this even for DPO in its original form. We provide a theoretical explanation for this regularization effect using ideal DPO policy theory. We further show performance improvements over DPO in translation and mathematical reasoning using our DPO variants. We find it can be beneficial to include ties in preference optimization rather than simply discard them, as is done in common practice.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- COOPERA: Continual Open-Ended Human-Robot AssistanceChenyang Ma, Kai Lu, Ruta Desai, Xavier Puig 等NeurIPS 2025 · 被引用 9 次
- Learning Guarantee of Reward Modeling Using Deep Neural NetworksYuanhang Luo, Yeheng Ge, Ruijian Han, Guohao ShenKDD 2026 · 被引用 3 次
- Reward Modeling with Ordinal Feedback: Wisdom of the CrowdShang Liu, Yu Pan, Guanting Chen, Xiaocheng LiICML 2025
- DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and Its Loss' Convexity is Dispensable)Wenxuan Zhou, Shujian Zhang, brice magdalou, John Lambert 等ICML 2026
它引用的顶会 Paper14
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
- SimPO: Simple Preference Optimization with a Reference-Free RewardYu Meng, Mengzhou Xia, Danqi ChenNeurIPS 2024 · 被引用 1,203 次
- Model Alignment as Prospect Theoretic OptimizationKawin Ethayarajh, Winnie Xu, Niklas Muennighoff, Dan Jurafsky 等ICML 2024 · 被引用 973 次
- Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine TranslationHaoran Xu, Amr Sharaf, Yunmo Chen, Weiting Tan 等ICML 2024 · 被引用 447 次
- Statistical Rejection Sampling Improves Preference OptimizationTianqi Liu, Yao Zhao, Rishabh Joshi, Misha Khalman 等ICLR 2024 · 被引用 346 次
相关 Paper
- TokenRatio: Principled Token-Level Preference Optimization via Ratio MatchingTruong Nguyen, Tien-Phat Nguyen, Linh Van, Duy Nguyen 等ICML 2026
- TODO: Enhancing LLM Alignment with Ternary PreferencesYuxiang Guo, Lu Yin, Bo Jiang, Jiaqi ZhangICLR 2025
- Uncertainty-Aware Iterative Preference Optimization for Enhanced LLM ReasoningLei Li, Hehuan Liu, Yaxin Zhou, ZhaoYang Gui 等ACL 2025 · 被引用 3 次
- Reward Alignment Optimization: A Direct Point-wise Alignment ApproachZelin Li, Jia Leng, Dawei Song, Yangen HuACL 2026
- Spurious Correlation Learning in Preference Optimization: Mechanisms, Consequences, and Mitigation via Tie TrainingChristian Moya, Alex Semendinger, Guang Lin, Elliott ThornleyICML 2026 · 被引用 1 次
