Quagmires in SFT-RL Post-Training: When High SFT Scores Mislead and What to Use Instead
Feiyang Kang, Michael Kuchnik, Karthik Padthe, Marin Vlastelica, Ruoxi Jia, Carole-Jean Wu, Newsha Ardalani
摘要
In post-training for reasoning Large Language Models (LLMs), the current state of practice trains LLMs in two independent stages: Supervised Fine-Tuning (SFT) and Reinforcement Learning with Verifiable Rewards (RLVR, shortened as "RL" below). In this work, we challenge whether high SFT scores translate to improved performance after RL. We provide extensive counter-examples where this is not true. We find high SFT scores can be biased toward simpler or more homogeneous data and are not reliably predictive of subsequent RL gains or scaled-up post-training effectiveness. In some cases, RL training on models with improved SFT performance could lead to substantially worse outcome compared to RL on the base model without SFT. We study alternative metrics and identify generalization loss on held-out reasoning examples and Pass@large k performance to provide strong proxies for the RL outcome. We trained hundreds of models up to 12B-parameter with SFT and RLVR via GRPO and ran extensive evaluations on 7 math benchmarks with up to 256 repetitions, spending 1M GPU hours. Experiments include models from Llama3, Mistral-Nemo, Qwen3 and multiple state-of-the-art SFT/RL datasets. Compared to directly predicting from pre-RL performance, prediction based on generalization loss and Pass@large k achieves substantial higher precision, improving coefficient and Spearman's rank correlation coefficient by up to 0.5 (2x). This provides strong utility for broad use cases. For example, in most experiments, we find SFT training on unique examples for a one epoch underperforms training on half examples for two epochs, either after SFT or SFT-then-RL; With the same SFT budget, training only on short examples may lead to better SFT performance, though, it often leads to worse outcome after RL compared to training on examples with varying lengths. This work develops an enhanced evaluation tool for math reasoning tasks and is open-sourced.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Good SFT Optimizes for SFT, Better SFT Prepares for Reinforcement LearningDylan Zhang, Yufeng Xu, Haojin Wang, Qingzhi Chen 等ICML 2026 · 被引用 8 次
- Learning to Edit Knowledge via Instruction-based Chain-of-Thought PromptingJinhu Fu, Yan Bai, Longzhu He, Yihang Lou 等ACL 2026 · 被引用 1 次
- Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic OptimizationYihang Yao, Zhepeng Cen, Haohong Lin, Shiqi Liu 等ICML 2026
- STAR: Similarity-guided Teacher-Assisted Refinement for Super-Tiny Function Calling ModelsJiliang Ni, Jiachen Pu, Zhongyi Yang, Jingfeng Luo 等ICLR 2026
- Rethinking Expert Trajectory Utilization in LLM Post-training for Mathematical ReasoningBowen Ding, Yuhan Chen, Jiayang Lyu, Jiyao Yuan 等ACL 2026
它引用的顶会 Paper14
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer 等NeurIPS 2022 · 被引用 2,039 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM ReasoningShenzhi Wang, Le Yu, Chang Gao, Chujie Zheng 等NeurIPS 2025 · 被引用 592 次
- ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language ModelsMingjie Liu, Shizhe Diao, Ximing Lu, Jian Hu 等NeurIPS 2025 · 被引用 181 次
- AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement LearningYang Chen, Zhuolin Yang, Zihan Liu, Chankyu Lee 等NeurIPS 2025 · 被引用 79 次
相关 Paper
- Reinforcement Learning for Reasoning in Large Language Models with One Training ExampleYiping Wang, Qing Yang, Zhiyuan Zeng, Liliang Ren 等NeurIPS 2025 · 被引用 314 次
- Breaking Barriers: Do Reinforcement Post Training Gains Transfer To Unseen Domains?Chuxuan Hu, Yuxuan Zhu, Antony Kellermann, Caleb Biddulph 等ICLR 2026
- Spurious Rewards: Rethinking Training Signals in RLVRRulin Shao, Stella Li, Rui Xin, Scott Geng 等ICML 2026
- Beyond Two-Stage Training: Cooperative SFT and RL for LLM ReasoningLiang Chen, Xueting Han, Li Shen, Jing Bai 等ICML 2026 · 被引用 24 次
- First SFT, Second RL, Third UPT: Continual Improving Multi-Modal LLM Reasoning via Unsupervised Post-TrainingLai Wei, Yuting Li, Chen Wang, Yue Wang 等NeurIPS 2025 · 被引用 28 次
