Lune

NeurIPS2025Top-tier venue

Preference Distillation via Value based Reinforcement Learning

Minchan Kwon, Junwon Ko, Kangil Kim, Junmo Kim

2025Year
1Top-tier citations

Abstract

Direct Preference Optimization (DPO) is a powerful paradigm to align language models with human preferences using pairwise comparisons. However, its binary win-or-loss supervision often proves insufficient for training small models with limited capacity. Prior works attempt to distill information from large teacher models using behavior cloning or KL divergence. These methods often focus on mimicking current behavior and overlook distilling reward modeling. To address this issue, we propose Teacher Value-based Knowledge Distillation (TVKD), which introduces an auxiliary reward from the value function of the teacher model to provide a soft guide. This auxiliary reward is formulated to satisfy potential-based reward shaping, ensuring that the global reward structure and optimal policy of DPO are preserved. TVKD can be integrated into the standard DPO training framework and does not require additional rollouts. Our experimental results show that TVKD consistently improves performance across various benchmarks and model sizes. * Co-supervising authors. 39th Conference on Neural Information Processing Systems (NeurIPS 2025).

teacher in the RL objective may result in conflict between following the teacher policy and optimizing for human preferences in the offline dataset. This undermines the original intent of DPO, which is to align policies with human preferences expressed in data.

In this work, we propose Teacher Value-based Knowledge Distillation (TVKD), a method that integrates soft reward labels from the teacher without interfering with the DPO reward structure and its global optimal policy. The key idea is to introduce a novel auxiliary reward term that leverages the value function of the teacher. These value functions provide estimates of the value function that serves as a proxy for the internal reward modeling of the teacher model. We add this value function to the RL objective in a form that satisfies Potential-based Reward shaping (PBRS) [20]. This formula theoretically guarantees that the comparison of the action-level Q-functions will not change, maintaining the optimal policy. We introduce these value functions as a soft sequence-level reward, enabling DPO to account not just for which sequence wins, but by how much. This adds more informative supervision to the originally binary feedback. Practically, it can be integrated into the DPO loss with minimal modification, providing a simple yet effective way to incorporate internal signals from the teacher model.

We demonstrate that TVKD achieves strong performance across multiple benchmarks for preference distillation, including AlpacaEval [6], and the Open LLM Leaderboard [8]. Notably, our method requires no additional rollouts and only leverages teacher outputs on an existing offline DPO dataset, making it compatible with the standard training frameworks of DPO. In addition, we demonstrate that TVKD performs consistently across student models ranging from 0.5B to 3B parameters, and remains robust under a wide range of ablation.

Our main contributions are summarized as follows:

• We introduce Teacher Value-based Knowledge Distillation, a preference distillation method that leverages the value function of the teacher as a proxy for its reward model.

• We present a form of auxiliary reward term that satisfies potential-based reward shaping, suggesting a method for adding teacher information without conflict with the original DPO reward structure.

• We empirically validate TVKD on multiple preference benchmarks, demonstrating consistent improvements across various student model sizes and datasets.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 9dbcf3b3-d161-47ef-9c51-c07592548b1f

Cited by top-tier papers1

Ask how each one uses it

Builds on9

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines