Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling
Zeyu Huang, Tianhao Cheng, Zihan Qiu, Zili Wang, Xu Yinghui, Edoardo Ponti, Ivan Titov
Abstract
Existing LLMs-post-training techniques are broadly categorized into supervised fine-tuning (SFT) and reinforcement fine-tuning (RFT). Each paradigm presents a distinct trade-off: (1) SFT excels at mimicking demonstration data, but can lead to problematic generalization as a form of behavior cloning. (2) Conversely, RFT can significantly enhance a model's performance but is prone to learning unexpected behaviors, and its performance is sensitive to the initial policy. In this paper, we propose a unified view of these methods and introduce Prefix-RFT, a hybrid approach that synergizes learning from both demonstration and exploration. Using mathematical reasoning problems as a test bed, we empirically demonstrate that Prefix-RFT is simple yet effective. Not only does it surpass the performance of standalone SFT and RFT, but it also outperforms parallel mixed-policy RFT methods. Our analysis highlights the complementary nature of SFT and RFT, validating that Prefix-RFT effectively harmonizes them. Further ablation studies confirm the method's robustness to variations in the quality and quantity of demonstration data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7d584fa3-618d-4483-b90b-5131a5f8dbcdCited by top-tier papers10
- On-Policy RL Meets Off-Policy Experts: Harmonizing Supervised Fine-Tuning and Reinforcement Learning via Dynamic WeightingWenhao Zhang, Yuexiang Xie, Yuchang Sun, Yanxi Chen et al.ICLR 2026 · 100 citations
- Scaf-GRPO: Scaffolded Group Relative Policy Optimization for Enhancing LLM ReasoningXichen Zhang, Sitong Wu, Yinghao Zhu, Haoru Tan et al.ICLR 2026 · 52 citations
- Inpainting-Guided Policy Optimization for Diffusion Large Language ModelsSiyan Zhao, Mengchen Liu, Jing Huang, Miao Liu et al.ICLR 2026 · 14 citations
- Reuse your FLOPs: Scaling RL on Hard Problems by Conditioning on Very Off-Policy PrefixesAmrith Setlur, Zijian Wang, Andrew Cohen, Paria Rashidinejad et al.ICML 2026 · 13 citations
- SABER: Switchable and Balanced Training for Efficient LLM ReasoningKai Zhao, Yanjun Zhao, Jiaming Song, Shien He et al.AAAI 2026 · 9 citations
Builds on14
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer et al.NeurIPS 2022 · 2,039 citations
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?Zhiqi Chen, Rui Lu, Andrew Zhao, Zhaokai Wang et al.NeurIPS 2025 · 1,109 citations
- Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base ModelJingcheng Hu, Yinmin Zhang, Qi Han, Daxin Jiang et al.NeurIPS 2025 · 533 citations
- Efficient Online Reinforcement Learning with Offline DataPhilip J. Ball, Laura Smith, Ilya Kostrikov, Sergey LevineICML 2023 · 326 citations
Related papers
- UFT: Unifying Supervised and Reinforcement Fine-TuningMingyang Liu, Gabriele Farina, Asuman OzdaglarNeurIPS 2025 · 61 citations
- Trust-Region Adaptive Policy OptimizationMingyu Su, Jian Guan, Yuxian Gu, Minlie Huang et al.ICLR 2026 · 2 citations
- Retaining by Doing: The Role of On-Policy Data in Mitigating ForgettingHoward Chen, Noam Razin, Karthik Narasimhan, Danqi ChenICML 2026
- Learning What Reinforcement Learning Can't: Interleaved Online Fine-Tuning for Hardest QuestionsLu Ma, Hao Liang, Meiyi Qiang, Lexiang Tang et al.ICLR 2026 · 103 citations
- Learning Diverse Responses with Prefix-Conditioned Supervised Fine-TuningZhiyuan Fan, Guanqiao Chen, Yanyi Huang, Mingkuan Zhao et al.ACL 2026
