Skip-Thinking: Chunk-wise Chain-of-Thought Distillation Enable Smaller Language Models to Reason Better and Faster
Xiaoshu Chen, Sihang Zhou, Ke Liang, Xiaoyu Sun, Xinwang Liu
摘要
Chain-of-thought (CoT) distillation allows a large language model (LLM) to guide a small language model (SLM) in reasoning tasks. Existing methods train the SLM to learn the long rationale in one iteration, resulting in two issues: 1) Long rationales lead to a large tokenlevel batch size during training, making gradients of core reasoning tokens (i.e., the token will directly affect the correctness of subsequent reasoning) over-smoothed as they contribute a tiny fraction of the rationale. As a result, the SLM converges to sharp minima where it fails to grasp the reasoning logic. 2) The response is slow, as the SLM must generate a long rationale before reaching the answer. Therefore, we propose chunk-wise training (CWT), which uses a heuristic search to divide the rationale into internal semantically coherent chunks and focuses SLM on learning from only one chunk per iteration. In this way, CWT naturally isolates non-reasoning chunks that do not involve the core reasoning token (e.g., summary and transitional chunks) from the SLM learning for reasoning chunks, making the fraction of the core reasoning token increase in the corresponding iteration. Based on CWT, skip-thinking training (STT) is proposed. STT makes the SLM automatically skip nonreasoning medium chunks to reach the answer, improving reasoning speed while maintaining accuracy. We validate our approach on a variety of SLMs and multiple reasoning tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- DenseSteer: Steering Small Language Models towards Dense Math ReasoningYang Ouyang, Shuhang Lin, Jung-Eun KimICML 2026 · 被引用 1 次
- ImgCoT: Compressing Long Chain of Thought into Compact Visual Tokens for Efficient Reasoning of Large Language ModelXiaoshu Chen, sihang zhou, KE LIANG, Taichun Zhou 等ICML 2026 · 被引用 1 次
- Are Rationales Necessary and Sufficient? Tuning LLMs for Explainable Misinformation DetectionBing Wang, Rui Miao, Ximing Li, Chen Shen 等KDD 2026 · 被引用 1 次
- Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative ExplorationShuzhang Zhong, Haochen Huang, Shengxuan Qiu, Pengfei Zuo 等OSDI 2026
它引用的顶会 Paper7
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Think before you speak: Training Language Models With Pause TokensSachin Goyal, Ziwei Ji, Ankit Singh Rawat, Aditya Krishna Menon 等ICLR 2024 · 被引用 240 次
- Large Language Models Are Reasoning TeachersNamgyu Ho, Laura Schmid, Se-Young YunACL 2023 · 被引用 102 次
相关 Paper
- Investigating Mysteries of CoT-Augmented DistillationSomin Wadhwa, Silvio Amir, Byron C. WallaceEMNLP 2024 · 被引用 1 次
- UniCoTT: A Unified Framework for Structural Chain-of-Thought DistillationXianwei Zhuang, Zhihong Zhu, Zhichang Wang, Xuxin Cheng 等ICLR 2025
- CODI: Compressing Chain-of-Thought into Continuous Space via Self-DistillationZhenyi Shen, Hanqi Yan, Linhai Zhang, Zhanghao Hu 等EMNLP 2025
- Think, But Don't Tell: Implicit Reasoning for LLM-based Sequential Recommendation via Multi-Teacher DistillationWeihai Lu, Xiaoxi Cui, Chenke YinSIGIR 2026
- Improving Reasoning Capabilities in Small Models through Mixture-of-layers Distillation with Stepwise Attention on Key InformationYao Chen, Jiawei Sheng, Wenyuan Zhang, Tingwen LiuEMNLP 2025
