KL Penalty Control via Perturbation for Direct Preference Optimization
Sangkyu Lee, Janghoon Han, Hosung Song, Stanley Jungkyu Choi, Honglak Lee, Youngjae Yu
摘要
Direct Preference Optimization (DPO) demonstrates the advantage of aligning a large language model with human preference using only an offline dataset. However, DPO has the limitation that the KL penalty, which prevents excessive deviation from the reference model, is static throughout the training process. Several methods claim to change this static KL penalty of DPO into a dynamic one, but no approach can adaptively assign different KL penalties for each preference pair. In this paper, we propose -Direct Preference Optimization (-DPO), which allows adaptive control of the KL penalty strength for each preference pair. Specifically, -DPO adaptively controls for each preference pair based on the monotonicity of logits as a preference model under the perturbation of during training. This is equivalent to adjusting the KL penalty by checking whether the change in training-time temperature can lead to better preference confidence as preference models by simply reusing the logit of the current policy and the reference policy. Experimental results show that the simple criterion of -DPO for KL penalty relaxation significantly improves DPO compared to most existing direct alignment algorithms on general chatbot benchmarks and reveal that this KL penalty control criterion can reflect confusion as a preference model and provide an efficient KL trade-off, highlighting the significance of instance-level adaptive KL penalty control in DPO.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- -DPO: Robust Preference Alignment for Diffusion Models via DivergenceYang Li, Songlin Yang, Wei Wang, Xiaoxuan Han 等ICLR 2026
- Displacement-Resistant Extensions of DPO with Nonconvex -DivergencesIdan Pipano, Shoham Sabach, Kavosh Asadi, Mohammad GhavamzadehICLR 2026
它引用的顶会 Paper19
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 被引用 3,228 次
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- SimPO: Simple Preference Optimization with a Reference-Free RewardYu Meng, Mengzhou Xia, Danqi ChenNeurIPS 2024 · 被引用 1,203 次
相关 Paper
- Cal-DPO: Calibrated Direct Preference Optimization for Language Model AlignmentTeng Xiao, Yige Yuan, Huaisheng Zhu, Mingxiao Li 等NeurIPS 2024 · 被引用 76 次
- ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference OptimizationHee Suk Yoon, Eunseop Yoon, Mark A. Hasegawa-Johnson, Sungwoong Kim 等ICML 2025
- AlphaDPO: Adaptive Reward Margin for Direct Preference OptimizationJunkang Wu, Xue Wang, Zhengyi Yang, Jiancan Wu 等ICML 2025
- Private Direct Preference Optimization for LLM AlignmentYangfan Jiang, Fei Wei, Ergute Bao, Xiaokui Xiao 等CCS 2026
- Keep the Best, Forget the Rest: Reliable Alignment with Order-Aware Preference OptimizationJiahui Zhu, Yuanjie Shi, Xiyue Peng, Xin Liu 等ICLR 2026
