Improving Reasoning Capabilities in Small Models through Mixture-of-layers Distillation with Stepwise Attention on Key Information
Yao Chen, Jiawei Sheng, Wenyuan Zhang, Tingwen Liu
摘要
The significant computational demands of large language models have increased interest in distilling reasoning abilities into smaller models via Chain-of-Thought (CoT) distillation. Current CoT distillation methods mainly focus on transferring teacher-generated rationales for complex reasoning to student models. However, they do not adequately explore teachers' dynamic attention toward critical information during reasoning. We find that language models exhibit progressive attention shifts towards key information during reasoning, which implies essential clues for drawing conclusions. Building on this observation and analysis, we introduce a novel CoT distillation framework that transfers the teacher's stepwise attention on key information to the student model. This establishes structured guidance for the student's progressive concentration on key information during reasoning. More importantly, we develop a Mixture of Layers module enabling dynamic alignment that adapts to different layers between the teacher and student. Our method achieves consistent performance improvements across multiple mathematical and commonsense reasoning datasets. To our knowledge, it is the first method to leverage stepwise attention within CoT distillation to improve small model reasoning. Question: (Step1)A mailman has to give 4 pieces of junk mail to each house in each of the 16 blocks. If there are 17 houses in each block, How many pieces of junk mail should he give in total? Rationale: (Step2)The mailman has to give 4 pieces of junk mail to each house in each of the 16 blocks. (Step3)There are 17 houses in each block. (Step4)So, the total number of houses is 16 * 17 = 272 houses. (Step5)The mailman has to give 4 pieces of junk mail to each house, so the total number of junk mail pieces is 272 * 4 = 1088. Answer: (16 * 17 * 4 ) Question Answer
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Logical Phase Transitions: Understanding Collapse in LLM Logical ReasoningXinglang Zhang, Yunyao Zhang, ZeLiang Chen, Junqing Yu 等ACL 2026 · 被引用 28 次
- STRIDE-ED: A Strategy-Grounded Stepwise Reasoning Framework for Empathetic Dialogue SystemsHongru Ji, Yuyin Fan, Meng Zhao, Xianghua Li 等ACL 2026 · 被引用 1 次
它引用的顶会 Paper16
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersWenhui Wang, Furu Wei, Li Dong, Hangbo Bao 等NeurIPS 2020 · 被引用 2,727 次
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
- Mixture-of-Experts with Expert Choice RoutingYanqi Zhou, Tao Lei, Hanxiao Liu, Nan Du 等NeurIPS 2022 · 被引用 933 次
相关 Paper
- Investigating Mysteries of CoT-Augmented DistillationSomin Wadhwa, Silvio Amir, Byron C. WallaceEMNLP 2024 · 被引用 1 次
- Mentor-KD: Making Small Language Models Better Multi-step ReasonersHojae Lee, Junho Kim, SangKeun LeeEMNLP 2024
- Improve Student's Reasoning Generalizability through Cascading Decomposed CoTs DistillationChengwei Dai, Kun Li, Wei Zhou, Songlin HuEMNLP 2024 · 被引用 4 次
- Long-Chain Reasoning Distillation via Adaptive Prefix AlignmentZhenghao Liu, Zhuoyang Wu, Xinze Li, Yukun Yan 等ACL 2026 · 被引用 4 次
- Capture the Key in Reasoning to Enhance CoT Distillation GeneralizationChengwei Dai, Kun Li, Wei Zhou, Songlin HuACL 2025
