NFT: Bridging Supervised Learning and Reinforcement Learning in Math Reasoning
Huayu Chen, Kaiwen Zheng, Qinsheng Zhang, Ganqu Cui, Yin Cui, Haotian Ye, Tsung-Yi Lin, Ming-Yu Liu, Jun Zhu, Haoxiang Wang
摘要
Reinforcement Learning (RL) has played a central role in the recent surge of LLMs' math abilities by enabling verification-driven training through binary verifier signals. In contrast, Supervised Learning (SL) is rarely considered for such verification-driven training, largely due to its heavy reliance on reference answers and inability to reflect on mistakes. In this work, we challenge the prevailing notion that self-improvement is exclusive to RL and propose Negative-aware Fine-Tuning (NFT) --- a supervised approach that enables LLMs to reflect on their failures and improve autonomously with no external teachers. In online training, instead of throwing away self-generated negative answers, NFT constructs an implicit negative policy to model them. This implicit policy is parameterized with the same positive LLM we target to optimize on positive data, enabling direct policy optimization on all LLMs' generations. We conduct experiments on 7B and 32B models in math reasoning tasks. Results consistently show that through the additional leverage of negative feedback, NFT significantly improves over SL baselines like rejection fine-tuning, matching, or even surpassing leading RL algorithms like GRPO and DAPO. Furthermore, we demonstrate that NFT and GRPO are actually equivalent in strict-on-policy training, even though they have entirely different theoretical foundations. Our experiments and theoretical findings bridge the gap between SL and RL methods in binary-feedback learning systems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Learning to Reason under Off-Policy GuidanceJianhao Yan, Yafu Li, Zican Hu, Zhi Wang 等NeurIPS 2025 · 被引用 310 次
- Trajectory Bellman Residual Minimization: A Simple Value-Based Method for LLM ReasoningYurun Yuan, Fan Chen, Zeyu Jia, Alexander Rakhlin 等NeurIPS 2025 · 被引用 10 次
- EvoID: Reinforced Evolution for Identity-Preserving Video GenerationYiheng Zhang, Zhaofan Qiu, Zunxu Liu, Yingwei Pan 等CVPR 2026
- Compatibility-Aware Dynamic Fine-Tuning for Large Language ModelsYucheng Zhou, Junwei Sheng, Qianning Wang, Jianbing ShenACL 2026
- Mimic Human Cognition, Master Multi-Image Reasoning: A Meta-Action Framework for Enhanced Visual UnderstandingJianghao Yin, Qingbin Li, Kun Sun, Cheng Ding 等CVPR 2026
相关 Paper
- Quagmires in SFT-RL Post-Training: When High SFT Scores Mislead and What to Use InsteadFeiyang Kang, Michael Kuchnik, Karthik Padthe, Marin Vlastelica 等ICLR 2026 · 被引用 27 次
- ReFT: Reasoning with Reinforced Fine-TuningLuong Quoc Trung, Xinbo Zhang, Zhanming Jie, Peng Sun 等ACL 2024
- Native Reasoning Models: Training Language Models to Reason on Unverifiable DataYuanfu Wang, Zhixuan Liu, Li xiangtian, Chaochao Lu 等ICLR 2026 · 被引用 3 次
- Continual SFT Matches Multimodal RLHF with Negative SupervisionKe Zhu, Yu Wang, Yanpeng Sun, Qiang Chen 等CVPR 2025
- DiffusionNFT: Online Diffusion Reinforcement with Forward ProcessKaiwen Zheng, Huayu Chen, Haotian Ye, Haoxiang Wang 等ICLR 2026 · 被引用 213 次
