Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment
Teng Xiao, Yige Yuan, Huaisheng Zhu, Mingxiao Li, Vasant G. Honavar
Abstract
We study the problem of aligning large language models (LLMs) with human preference data. Contrastive preference optimization has shown promising results in aligning LLMs with available preference data by optimizing the implicit reward associated with the policy. However, the contrastive objective focuses mainly on the relative values of implicit rewards associated with two responses while ignoring their actual values, resulting in suboptimal alignment with human preferences. To address this limitation, we propose calibrated direct preference optimization (Cal-DPO), a simple yet effective algorithm. We show that substantial improvement in alignment with the given preferences can be achieved simply by calibrating the implicit reward to ensure that the learned implicit rewards are comparable in scale to the ground-truth rewards. We demonstrate the theoretical advantages of Cal-DPO over existing approaches. The results of our experiments on a variety of standard benchmarks show that Cal-DPO remarkably improves off-the-shelf methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aba70552-77a6-4a14-bc8b-a1a5e2bd4ca7Cited by top-tier papers19
- ComPO: Preference Alignment via Comparison OraclesPeter Chen, Xi Chen, Wotao Yin, Tianyi LinNeurIPS 2025 · 20 citations
- Inference-time Alignment in Continuous SpaceYige Yuan, Teng Xiao, Yunfan Li, Bingbing Xu et al.NeurIPS 2025 · 9 citations
- Proximalized Preference Optimization for Diverse Feedback Types: A Decomposed Perspective on DPOKaiyang Guo, Yinchuan Li, Zhitang ChenNeurIPS 2025 · 7 citations
- SPACE: Noise Contrastive Estimation Stabilizes Self-Play Fine-Tuning for Large Language ModelsYibo Wang, Guangda Huzhang, Qingguo Chen, Zhao Xu et al.NeurIPS 2025 · 7 citations
- Simple Distillation for One-Step Diffusion ModelsHuaisheng Zhu, Teng Xiao, Shijie Zhou, Zhimeng Guo et al.NeurIPS 2025 · 7 citations
Builds on24
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- Self-Rewarding Language ModelsWeizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Xian Li et al.ICML 2024 · 569 citations
- Self-Play Fine-Tuning Converts Weak Language Models to Strong Language ModelsZixiang Chen, Yihe Deng, Huizhuo Yuan, Kaixuan Ji et al.ICML 2024 · 527 citations
Related papers
- What Matters in Data for DPO?Yu Pan, Zhongze Cai, Huaiyang Zhong, Guanting Chen et al.NeurIPS 2025 · 13 citations
- Private Direct Preference Optimization for LLM AlignmentYangfan Jiang, Fei Wei, Ergute Bao, Xiaokui Xiao et al.CCS 2026
- mDPO: Conditional Preference Optimization for Multimodal Large Language ModelsFei Wang, Wenxuan Zhou, James Y. Huang, Nan Xu et al.EMNLP 2024 · 11 citations
- Towards Efficient Exact Optimization of Language Model AlignmentHaozhe Ji, Cheng Lu, Yilin Niu, Pei Ke et al.ICML 2024 · 32 citations
- Direct Large Language Model Alignment Through Self-Rewarding Contrastive Prompt DistillationAiwei Liu, Haoping Bai, Zhiyun Lu, Xiang Kong et al.ACL 2024 · 4 citations
