Lune

NeurIPS2025顶会

Preference Distillation via Value based Reinforcement Learning

Minchan Kwon, Junwon Ko, Kangil Kim, Junmo Kim

2025年份
1顶会引用

摘要

Direct Preference Optimization (DPO) is a powerful paradigm to align language models with human preferences using pairwise comparisons. However, its binary win-or-loss supervision often proves insufficient for training small models with limited capacity. Prior works attempt to distill information from large teacher models using behavior cloning or KL divergence. These methods often focus on mimicking current behavior and overlook distilling reward modeling. To address this issue, we propose Teacher Value-based Knowledge Distillation (TVKD), which introduces an auxiliary reward from the value function of the teacher model to provide a soft guide. This auxiliary reward is formulated to satisfy potential-based reward shaping, ensuring that the global reward structure and optimal policy of DPO are preserved. TVKD can be integrated into the standard DPO training framework and does not require additional rollouts. Our experimental results show that TVKD consistently improves performance across various benchmarks and model sizes. * Co-supervising authors. 39th Conference on Neural Information Processing Systems (NeurIPS 2025).

teacher in the RL objective may result in conflict between following the teacher policy and optimizing for human preferences in the offline dataset. This undermines the original intent of DPO, which is to align policies with human preferences expressed in data.

In this work, we propose Teacher Value-based Knowledge Distillation (TVKD), a method that integrates soft reward labels from the teacher without interfering with the DPO reward structure and its global optimal policy. The key idea is to introduce a novel auxiliary reward term that leverages the value function of the teacher. These value functions provide estimates of the value function that serves as a proxy for the internal reward modeling of the teacher model. We add this value function to the RL objective in a form that satisfies Potential-based Reward shaping (PBRS) [20]. This formula theoretically guarantees that the comparison of the action-level Q-functions will not change, maintaining the optimal policy. We introduce these value functions as a soft sequence-level reward, enabling DPO to account not just for which sequence wins, but by how much. This adds more informative supervision to the originally binary feedback. Practically, it can be integrated into the DPO loss with minimal modification, providing a simple yet effective way to incorporate internal signals from the teacher model.

We demonstrate that TVKD achieves strong performance across multiple benchmarks for preference distillation, including AlpacaEval [6], and the Open LLM Leaderboard [8]. Notably, our method requires no additional rollouts and only leverages teacher outputs on an existing offline DPO dataset, making it compatible with the standard training frameworks of DPO. In addition, we demonstrate that TVKD performs consistently across student models ranging from 0.5B to 3B parameters, and remains robust under a wide range of ablation.

Our main contributions are summarized as follows:

• We introduce Teacher Value-based Knowledge Distillation, a preference distillation method that leverages the value function of the teacher as a proxy for its reward model.

• We present a form of auxiliary reward term that satisfies potential-based reward shaping, suggesting a method for adding teacher information without conflict with the original DPO reward structure.

• We empirically validate TVKD on multiple preference benchmarks, demonstrating consistent improvements across various student model sizes and datasets.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper1

问问它们各自怎么用它

它引用的顶会 Paper9

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖