Investigating Mysteries of CoT-Augmented Distillation
Somin Wadhwa, Silvio Amir, Byron C. Wallace
Abstract
Eliciting chain of thought (CoT) rationales - sequences of token that convey a “reasoning” process has been shown to consistently improve LLM performance on tasks like question answering. More recent efforts have shown that such rationales can also be used for model distillation: Including CoT sequences (elicited from a large “teacher” model) in addition to target labels when fine-tuning a small student model yields (often substantial) improvements. In this work we ask: Why and how does this additional training signal help in model distillation? We perform ablations to interrogate this, and report some potentially surprising results. Specifically: (1) Placing CoT sequences after labels (rather than before) realizes consistently better downstream performance – this means that no student “reasoning” is necessary at test time to realize gains. (2) When rationales are appended in this way, they need not be coherent reasoning sequences to yield improvements; performance increases are robust to permutations of CoT tokens, for example. In fact, (3) a small number of key tokens are sufficient to achieve improvements equivalent to those observed when full rationales are used in model distillation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on8
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Specializing Smaller Language Models towards Multi-Step ReasoningYao Fu, Hao Peng, Litu Ou, Ashish Sabharwal et al.ICML 2023 · 347 citations
- Think before you speak: Training Language Models With Pause TokensSachin Goyal, Ziwei Ji, Ankit Singh Rawat, Aditya Krishna Menon et al.ICLR 2024 · 240 citations
- Revisiting Relation Extraction in the era of Large Language ModelsSomin Wadhwa, Silvio Amir, Byron C. WallaceACL 2023 · 145 citations
- Large Language Models Are Reasoning TeachersNamgyu Ho, Laura Schmid, Se-Young YunACL 2023 · 102 citations
Related papers
- SCOTT: Self-Consistent Chain-of-Thought DistillationPeifeng Wang, Zhengyang Wang, Zheng Li, Yifan Gao et al.ACL 2023 · 39 citations
- Mentor-KD: Making Small Language Models Better Multi-step ReasonersHojae Lee, Junho Kim, SangKeun LeeEMNLP 2024
- Improve Student's Reasoning Generalizability through Cascading Decomposed CoTs DistillationChengwei Dai, Kun Li, Wei Zhou, Songlin HuEMNLP 2024 · 4 citations
- Improving Reasoning Capabilities in Small Models through Mixture-of-layers Distillation with Stepwise Attention on Key InformationYao Chen, Jiawei Sheng, Wenyuan Zhang, Tingwen LiuEMNLP 2025
- Fine-Tuning on Diverse Reasoning Chains Drives Within-Inference CoT Refinement in LLMsHaritz Puerto, Tilek Chubakov, Xiaodan Zhu, Harish Tayyar Madabushi et al.ACL 2025 · 13 citations
