Turning Dust into Gold: Distilling Complex Reasoning Capabilities from LLMs by Leveraging Negative Data
Yiwei Li, Peiwen Yuan, Shaoxiong Feng, Boyuan Pan, Bin Sun, Xinglin Wang, Heda Wang, Kan Li
Abstract
Large Language Models (LLMs) have performed well on various reasoning tasks, but their inaccessibility and numerous parameters hinder wide application in practice. One promising way is distilling the reasoning ability from LLMs to small models by the generated chain-of-thought reasoning paths. In some cases, however, LLMs may produce incorrect reasoning chains, especially when facing complex mathematical problems. Previous studies only transfer knowledge from positive samples and drop the synthesized data with wrong answers. In this work, we illustrate the merit of negative data and propose a model specialization framework to distill LLMs with negative samples besides positive ones. The framework consists of three progressive steps, covering from training to inference stages, to absorb knowledge from negative data. We conduct extensive experiments across arithmetic reasoning tasks to demonstrate the role of negative data in distillation from LLM 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- Escape Sky-high Cost: Early-stopping Self-Consistency for Multi-step ReasoningYiwei Li, Peiwen Yuan, Shaoxiong Feng, Boyuan Pan et al.ICLR 2024 · 101 citations
- Neural-Symbolic Collaborative Distillation: Advancing Small Language Models for Complex Reasoning TasksHuanxuan Liao, Shizhu He, Yao Xu, Yuanzhe Zhang et al.AAAI 2025 · 17 citations
- C3PO: Optimized Large Language Model Cascades with Probabilistic Cost Constraints for ReasoningAntonios Valkanas, Soumyasundar Pal, Pavel Rumiantsev, Yingxue Zhang et al.NeurIPS 2025 · 10 citations
- Reasoning-to-Defend: Safety-Aware Reasoning Can Defend Large Language Models from JailbreakingJunda Zhu, Lingyong Yan, Shuaiqiang Wang, Dawei Yin et al.EMNLP 2025 · 2 citations
- Retrieved In-Context Principles from Previous MistakesHao Sun, Yong Jiang, Bo Wang, Yingyan Hou et al.EMNLP 2024 · 1 citation
Builds on17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer et al.NeurIPS 2022 · 2,039 citations
- Neural Text Generation With Unlikelihood TrainingSean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan et al.ICLR 2020 · 683 citations
Related papers
- Teaching Small Language Models Reasoning through Counterfactual DistillationTao Feng, Yicheng Li, Chenglin Li, Hao Chen et al.EMNLP 2024 · 1 citation
- QCRD: Quality-guided Contrastive Rationale Distillation for Large Language ModelsWei Wang, Zhaowei Li, Qi Xu, Yiqing Cai et al.EMNLP 2025 · 1 citation
- Democratizing Reasoning Ability: Tailored Learning from Large Language ModelZhaoyang Wang, Shaohan Huang, Yuxuan Liu, Jiahai Wang et al.EMNLP 2023 · 8 citations
- Specializing Smaller Language Models towards Multi-Step ReasoningYao Fu, Hao Peng, Litu Ou, Ashish Sabharwal et al.ICML 2023 · 347 citations
- Teach Small Models to Reason by Curriculum DistillationWangyi Jiang, Yaojie Lu, Hongyu Lin, Xianpei Han et al.EMNLP 2025
