Improving Reasoning Capabilities in Small Models through Mixture-of-layers Distillation with Stepwise Attention on Key Information
Yao Chen, Jiawei Sheng, Wenyuan Zhang, Tingwen Liu
Abstract
The significant computational demands of large language models have increased interest in distilling reasoning abilities into smaller models via Chain-of-Thought (CoT) distillation. Current CoT distillation methods mainly focus on transferring teacher-generated rationales for complex reasoning to student models. However, they do not adequately explore teachers' dynamic attention toward critical information during reasoning. We find that language models exhibit progressive attention shifts towards key information during reasoning, which implies essential clues for drawing conclusions. Building on this observation and analysis, we introduce a novel CoT distillation framework that transfers the teacher's stepwise attention on key information to the student model. This establishes structured guidance for the student's progressive concentration on key information during reasoning. More importantly, we develop a Mixture of Layers module enabling dynamic alignment that adapts to different layers between the teacher and student. Our method achieves consistent performance improvements across multiple mathematical and commonsense reasoning datasets. To our knowledge, it is the first method to leverage stepwise attention within CoT distillation to improve small model reasoning. Question: (Step1)A mailman has to give 4 pieces of junk mail to each house in each of the 16 blocks. If there are 17 houses in each block, How many pieces of junk mail should he give in total? Rationale: (Step2)The mailman has to give 4 pieces of junk mail to each house in each of the 16 blocks. (Step3)There are 17 houses in each block. (Step4)So, the total number of houses is 16 * 17 = 272 houses. (Step5)The mailman has to give 4 pieces of junk mail to each house, so the total number of junk mail pieces is 272 * 4 = 1088. Answer: (16 * 17 * 4 ) Question Answer
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 502f3a02-3f69-49b9-bae1-1ec7869f5791Cited by top-tier papers2
- Logical Phase Transitions: Understanding Collapse in LLM Logical ReasoningXinglang Zhang, Yunyao Zhang, ZeLiang Chen, Junqing Yu et al.ACL 2026 · 28 citations
- STRIDE-ED: A Strategy-Grounded Stepwise Reasoning Framework for Empathetic Dialogue SystemsHongru Ji, Yuyin Fan, Meng Zhao, Xianghua Li et al.ACL 2026 · 1 citation
Builds on16
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersWenhui Wang, Furu Wei, Li Dong, Hangbo Bao et al.NeurIPS 2020 · 2,727 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- Mixture-of-Experts with Expert Choice RoutingYanqi Zhou, Tao Lei, Hanxiao Liu, Nan Du et al.NeurIPS 2022 · 933 citations
Related papers
- Investigating Mysteries of CoT-Augmented DistillationSomin Wadhwa, Silvio Amir, Byron C. WallaceEMNLP 2024 · 1 citation
- Mentor-KD: Making Small Language Models Better Multi-step ReasonersHojae Lee, Junho Kim, SangKeun LeeEMNLP 2024
- Improve Student's Reasoning Generalizability through Cascading Decomposed CoTs DistillationChengwei Dai, Kun Li, Wei Zhou, Songlin HuEMNLP 2024 · 4 citations
- Long-Chain Reasoning Distillation via Adaptive Prefix AlignmentZhenghao Liu, Zhuoyang Wu, Xinze Li, Yukun Yan et al.ACL 2026 · 4 citations
- Capture the Key in Reasoning to Enhance CoT Distillation GeneralizationChengwei Dai, Kun Li, Wei Zhou, Songlin HuACL 2025
