Harnessing Negative Signals: Reinforcement Distillation from Teacher Data for LLM Reasoning
Shuyao Xu, Cheng Peng, Jiangxuan Long, Weidi Xu, Wei Chu, Yuan Qi
摘要
Recent advances in model distillation show that data from advanced reasoning models can effectively train smaller student models. However, standard practices discard incorrect reasoning traces -- valuable, yet underutilized data. This paper addresses the critical question: How can both positive and negative distilled reasoning traces be effectively leveraged to maximize LLM reasoning performance in an offline setting? We employ a two-stage training recipe: first, Supervised Fine-Tuning (SFT) on positive traces, followed by a refinement stage using both positive and negative traces. We find that a simple REINFORCE-style objective, which we term the Reinforcement Distillation (REDI) objective, outperforms established preference optimization methods like DPO and SimPO in this distillation context. Our empirical evaluations demonstrate the effectiveness of this approach. Notably, our Qwen-REDI-1.5B model, trained on just 131k traces from the open Open-R1 dataset, achieves an 83.1% score on MATH-500. Its performance matches that of DeepSeek-R1-Distill-Qwen-1.5B, a model trained on 800k proprietary data. This result showcases the remarkable data efficiency of utilizing previously discarded negative traces.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- MARD: Module-Aware Reasoning Distillation for Language Models with Adaptive SupervisionWenqi Yang, Jianjun Li, Zhibo Zhang, Mingqian Ding 等ACL 2026
- From Imitation to Discrimination: Toward a Generalized Curriculum Advantage Mechanism Enhancing Cross-Domain Reasoning TasksChangpeng Yang, Jinyang Wu, Yuchen Liu, Shuai Zhang 等AAAI 2026
它引用的顶会 Paper10
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer 等NeurIPS 2022 · 被引用 2,039 次
- SimPO: Simple Preference Optimization with a Reference-Free RewardYu Meng, Mengzhou Xia, Danqi ChenNeurIPS 2024 · 被引用 1,203 次
- Learning to Reason under Off-Policy GuidanceJianhao Yan, Yafu Li, Zican Hu, Zhi Wang 等NeurIPS 2025 · 被引用 310 次
- Iterative Reasoning Preference OptimizationRichard Yuanzhe Pang, Weizhe Yuan, He He, Kyunghyun Cho 等NeurIPS 2024 · 被引用 287 次
相关 Paper
- The Quality-Utility Paradox: Why High-Reward Data Impairs Small Model Mathematical ReasoningHaolong Qian, Xianliang Yang, Ma yinuo, Lirong Che 等ICML 2026
- AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement LearningYang Chen, Zhuolin Yang, Zihan Liu, Chankyu Lee 等NeurIPS 2025 · 被引用 79 次
- Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language ModelsSiyan Zhao, Zhihui Xie, Mengchen Liu, Jing Huang 等ICML 2026 · 被引用 245 次
- Making Expert Reasoning Learnable with Self-DistillationEthan Mendes, Jungsoo Park, Alan RitterICML 2026 · 被引用 1 次
- Where Did This Sentence Come From? Tracing Provenance in LLM Reasoning DistillationKaiyuan Liu, Shaotian Yan, Rui Miao, Bing Wang 等ICLR 2026 · 被引用 7 次
