Improve Student's Reasoning Generalizability through Cascading Decomposed CoTs Distillation
Chengwei Dai, Kun Li, Wei Zhou, Songlin Hu
摘要
Large language models (LLMs) exhibit enhanced reasoning at larger scales, driving efforts to distill these capabilities into smaller models via teacher-student learning. Previous works simply fine-tune student models on teachers' generated Chain-of-Thoughts (CoTs) data. Although these methods enhance indomain (IND) reasoning performance, they struggle to generalize to out-of-domain (OOD) tasks. We believe that the widespread spurious correlations between questions and answers may lead the model to preset a specific answer which restricts the diversity and generalizability of its reasoning process. In this paper, we propose Cascading Decomposed CoTs Distillation (CasCoD) to address these issues by decomposing the traditional single-step learning process into two cascaded learning steps. Specifically, by restructuring the training objectives-removing the answer from outputs and concatenating the question with the rationale as input-CasCoD's two-step learning process ensures that students focus on learning rationales without interference from the preset answers, thus improving reasoning generalizability. Extensive experiments demonstrate the effectiveness of CasCoD on both IND and OOD benchmark reasoning datasets 1 . * Kun Li is the corresponding author. 1 Code can be found at https://github.com/C-W-D/ CasCoD (a) Answer SFT consistently outperform Std-CoT on OOD tasks. (b) A case of spurious correla on between ques ons and answers. Question: Why did someone bring a swimsuit to a ski resort? Options: (A) To swim in a heated pool. (B) To wear as an underlayer for warmth. (C) To use as a fashion statement. (D) To participate in a polar bear plunge event.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Neural-Symbolic Collaborative Distillation: Advancing Small Language Models for Complex Reasoning TasksHuanxuan Liao, Shizhu He, Yao Xu, Yuanzhe Zhang 等AAAI 2025 · 被引用 17 次
- Pedagogically-Inspired Data Synthesis for Language Model Knowledge DistillationBowei He, Yankai Chen, Xiaokun Zhang, Linghe Kong 等ICLR 2026 · 被引用 2 次
- MIND: From Passive Mimicry to Active Reasoning through Capability-Aware Multi-Perspective CoT DistillationJin Cui, Jiaqi Guo, Jiepeng Zhou, Ruixuan Yang 等ACL 2026 · 被引用 1 次
- Capture the Key in Reasoning to Enhance CoT Distillation GeneralizationChengwei Dai, Kun Li, Wei Zhou, Songlin HuACL 2025
- Latent-Guided Reasoning: Empowering Small LLMs with Large-Model ThinkingHanzhu Chen, Lin Yang, Jie Wang, Junhao Yan 等ICLR 2026
它引用的顶会 Paper10
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
- How Language Model Hallucinations Can SnowballMuru Zhang, Ofir Press, William Merrill, Alisa Liu 等ICML 2024 · 被引用 406 次
相关 Paper
- Mentor-KD: Making Small Language Models Better Multi-step ReasonersHojae Lee, Junho Kim, SangKeun LeeEMNLP 2024
- Investigating Mysteries of CoT-Augmented DistillationSomin Wadhwa, Silvio Amir, Byron C. WallaceEMNLP 2024 · 被引用 1 次
- CODI: Compressing Chain-of-Thought into Continuous Space via Self-DistillationZhenyi Shen, Hanqi Yan, Linhai Zhang, Zhanghao Hu 等EMNLP 2025
- Improving Reasoning Capabilities in Small Models through Mixture-of-layers Distillation with Stepwise Attention on Key InformationYao Chen, Jiawei Sheng, Wenyuan Zhang, Tingwen LiuEMNLP 2025
- Teaching Small Language Models Reasoning through Counterfactual DistillationTao Feng, Yicheng Li, Chenglin Li, Hao Chen 等EMNLP 2024 · 被引用 1 次
