Towards On-Policy SFT: Distribution Discriminant Theory and its Applications in LLM Training
Miaosen Zhang, Yishan Liu, Shuxia Lin, Qi Dai, Chong Luo, Baining Guo, Weihao Jiang, Peng Hou, Anxiang Zeng, Xu Yang, Xin Geng
摘要
Supervised fine-tuning (SFT) is computationally efficient but often yields inferior generalization compared to reinforcement learning (RL). This gap is primarily driven by RL's use of on-policy data. We propose a framework to bridge this chasm by enabling On-Policy SFT. We first present Distribution Discriminant Theory (DDT), which explains and quantifies the alignment between data and the model-induced distribution. Leveraging DDT, we introduce two complementary techniques: (i) In-Distribution Finetuning (IDFT), a loss-level method to enhance generalization ability of SFT, and (ii) Hinted Decoding, a data-level technique that can re-align the training corpus to the model's distribution. Extensive experiments demonstrate that our framework achieves generalization performance surpassing prominent offline RL algorithms, including DPO and SimPO, while maintaining the efficiency of an SFT pipeline. The proposed framework thus offers a practical alternative in domains where RL is infeasible. We open-source the code here: https://github.com/zhangmiaosen2000/Towards- On-Policy-SFT.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper22
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer 等NeurIPS 2022 · 被引用 2,039 次
相关 Paper
- Discriminative Finetuning of Generative Large Language Models without Reward Models and Human Preference DataSiqi Guo, Ilgee Hong, Vicente Balmaseda, Changlong Yu 等ICML 2025
- On the Generalization of SFT: A Reinforcement Learning Perspective with Reward RectificationYongliang Wu, Yizhou Zhou, Ziheng Zhou, Yingzhe Peng 等ICLR 2026 · 被引用 130 次
- Implicit Reward as the Bridge: A Unified View of SFT and DPO ConnectionsBo Wang, Qinyuan Cheng, Runyu Peng, Rong Bao 等NeurIPS 2025 · 被引用 23 次
- Optimizing DDPM Sampling with Shortcut Fine-TuningYing Fan, Kangwook LeeICML 2023 · 被引用 95 次
- Preference-Oriented Supervised Fine-Tuning: Favoring Target Model over Aligned Large Language ModelsYuchen Fan, Yuzhong Hong, Qiushi Wang, Junwei Bao 等AAAI 2025 · 被引用 7 次
