Mentor-KD: Making Small Language Models Better Multi-step Reasoners
Hojae Lee, Junho Kim, SangKeun Lee
Abstract
Large Language Models (LLMs) have displayed remarkable performances across various complex tasks by leveraging Chain-of-Thought (CoT) prompting. Recently, studies have proposed a Knowledge Distillation (KD) approach, reasoning distillation, which transfers such reasoning ability of LLMs through fine-tuning language models of multi-step rationales generated by LLM teachers. However, they have inadequately considered two challenges regarding insufficient distillation sets from the LLM teacher model, in terms of 1) data quality and 2) soft label provision. In this paper, we propose Mentor-KD, which effectively distills the multi-step reasoning capability of LLMs to smaller LMs while addressing the aforementioned challenges. Specifically, we exploit a mentor, intermediate-sized task-specific fine-tuned model, to augment additional CoT annotations and provide soft labels for the student model during reasoning distillation. We conduct extensive experiments and confirm Mentor-KD's effectiveness across various models and complex reasoning tasks 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f0e2ded7-5208-4cd4-8820-eb665816fbf5Cited by top-tier papers9
- Distilling LLM Agent into Small Models with Retrieval and Code ToolsMinki Kang, Jongwon Jeong, Seanie Lee, Jaewoong Cho et al.NeurIPS 2025 · 51 citations
- Reasoning Scaffolding: Distilling the Flow of Thought from LLMsXiangyu Wen, Junhua Huang, Zeju Li, Min Li et al.ICLR 2026 · 7 citations
- StepER: Step-wise Knowledge Distillation for Enhancing Reasoning Ability in Multi-Step Retrieval-Augmented Language ModelsKyumin Lee, Minjin Jeon, Sanghwan Jang, Hwanjo YuEMNLP 2025 · 1 citation
- Learning from Diverse Reasoning Paths with Routing and CollaborationZhenyu Lei, Zhen Tan, Song Wang, Yaochen Zhu et al.EMNLP 2025
- Improving Reasoning Capabilities in Small Models through Mixture-of-layers Distillation with Stepwise Attention on Key InformationYao Chen, Jiawei Sheng, Wenyuan Zhang, Tingwen LiuEMNLP 2025
Builds on22
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersWenhui Wang, Furu Wei, Li Dong, Hangbo Bao et al.NeurIPS 2020 · 2,727 citations
Related papers
- Teaching Small Language Models Reasoning through Counterfactual DistillationTao Feng, Yicheng Li, Chenglin Li, Hao Chen et al.EMNLP 2024 · 1 citation
- Improve Student's Reasoning Generalizability through Cascading Decomposed CoTs DistillationChengwei Dai, Kun Li, Wei Zhou, Songlin HuEMNLP 2024 · 4 citations
- Investigating Mysteries of CoT-Augmented DistillationSomin Wadhwa, Silvio Amir, Byron C. WallaceEMNLP 2024 · 1 citation
- Capture the Key in Reasoning to Enhance CoT Distillation GeneralizationChengwei Dai, Kun Li, Wei Zhou, Songlin HuACL 2025
- Think, But Don't Tell: Implicit Reasoning for LLM-based Sequential Recommendation via Multi-Teacher DistillationWeihai Lu, Xiaoxi Cui, Chenke YinSIGIR 2026
