The Hidden Link between RLHF and Contrastive Learning
Xufei Lv, Kehai Chen, Haoyuan Sun, Xuefeng Bai, Min zhang, Houde Liu
摘要
Alignment of large language models (LLMs) with human values has recently garnered significant attention, with prominent examples including the canonical yet costly Reinforcement Learning from Human Feedback (RLHF) and the simple Direct Preference Optimization (DPO). In this work, we demonstrate that both RLHF and DPO can be interpreted from the perspective of mutual information (MI) maximization, uncovering a profound connection to contrastive learning. Within this framework, both RLHF and DPO can be interpreted as methods that performing contrastive learning based on the positive and negative samples derived from base model, leveraging the Donsker-Varadhan (DV) lower bound on MI (equivalently, the MINE estimator). Such paradigm further reveals why RLHF cannot correctly update when the probability of sampling the correct answer from the base model is close to 0, which often leads to RLHF failing to incentivize reasoning capacities in LLMs beyond what is already present in the base model. Building on the perspective, we replace the DV/MINE bound with the Jensen-Shannon (JS) MI estimator and propose the Mutual Information Optimization (MIO). Comprehensive theoretical analysis and extensive empirical evaluations demonstrate that MIO mitigates the late-stage decline in chosenlikelihood observed in DPO, achieving competitive or superior performance across various challenging reasoning and mathematical benchmarks. The code is available at: https://anonymous.4open.science/r/MIO-63E6/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Multiplayer Nash Preference OptimizationFang Wu, Xu Huang, Weihao Xuan, Zhiwei Zhang 等ICLR 2026 · 被引用 8 次
- Maximizing Mutual Information Between Prompt and Response Improves LLM Performance With No Additional DataHyunji (Alex) Nam, Haoran Li, Natasha JaquesICML 2026
它引用的顶会 Paper17
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer 等NeurIPS 2022 · 被引用 2,039 次
- SimPO: Simple Preference Optimization with a Reference-Free RewardYu Meng, Mengzhou Xia, Danqi ChenNeurIPS 2024 · 被引用 1,203 次
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?Zhiqi Chen, Rui Lu, Andrew Zhao, Zhaokai Wang 等NeurIPS 2025 · 被引用 1,109 次
相关 Paper
- Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence ConstraintsChaoqi Wang, Yibo Jiang, Chenghao Yang, Han Liu 等ICLR 2024 · 被引用 173 次
- Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-constraintWei Xiong, Hanze Dong, Chenlu Ye, Ziqi Wang 等ICML 2024 · 被引用 346 次
- Cal-DPO: Calibrated Direct Preference Optimization for Language Model AlignmentTeng Xiao, Yige Yuan, Huaisheng Zhu, Mingxiao Li 等NeurIPS 2024 · 被引用 76 次
- On a Connection Between Imitation Learning and RLHFTeng Xiao, Yige Yuan, Mingxiao Li, Zhengyu Chen 等ICLR 2025
- As Simple as Fine-tuning: LLM Alignment via Bidirectional Negative Feedback LossXin Mao, Huimin Xu, Feng-Lin Li, Ziqi Jin 等ICLR 2025
