MARD: Module-Aware Reasoning Distillation for Language Models with Adaptive Supervision
Wenqi Yang, Jianjun Li, Zhibo Zhang, Mingqian Ding, Yushen Fang
摘要
Multi-step reasoning remains challenging for language models with limited capacity. While recent reasoning distillation approaches transfer chain-of-thought supervision from large teacher models, they typically apply uniform supervision across all Transformer components, overlooking the fact that different modules contribute unequally to reasoning. We propose Module-Aware Reasoning Distillation, a parameter-efficient framework that explicitly targets key Transformer components for effective reasoning transfer. Through systematic analysis, we identify the feed-forward network projections and the output projection of self-attention as primary bottlenecks for reasoning. Based on these findings, we introduce lightweight adapter modules at these components while freezing the backbone parameters, enabling focused and efficient distillation. Our approach adopts an offline distillation setting, where a strong teacher model provides reasoning trajectories in advance, and incorporates an adaptive supervision strategy that adjusts the strength of reasoning-related losses according to problem difficulty. Experiments on mathematical reasoning benchmarks demonstrate consistent improvements over strong baselines, and ablation studies confirm the importance of both module-aware placement and adaptive supervision. Our code is available at https: //github.com/wqyang24/MARD.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu 等ACL 2023 · 被引用 540 次
- Specializing Smaller Language Models towards Multi-Step ReasoningYao Fu, Hao Peng, Litu Ou, Ashish Sabharwal 等ICML 2023 · 被引用 347 次
- Hierarchical Graph Network for Multi-hop Question AnsweringYuwei Fang, Siqi Sun, Zhe Gan, Rohit Pillai 等EMNLP 2020 · 被引用 157 次
- Knowledge-Augmented Reasoning Distillation for Small Language Models in Knowledge-Intensive TasksMinki Kang, Seanie Lee, Jinheon Baek, Kenji Kawaguchi 等NeurIPS 2023 · 被引用 128 次
- GLoRe: When, Where, and How to Improve LLM Reasoning via Global and Local RefinementsAlexander Havrilla, Sharath Chandra Raparthy, Christoforos Nalmpantis, Jane Dwivedi-Yu 等ICML 2024 · 被引用 105 次
相关 Paper
- StepER: Step-wise Knowledge Distillation for Enhancing Reasoning Ability in Multi-Step Retrieval-Augmented Language ModelsKyumin Lee, Minjin Jeon, Sanghwan Jang, Hwanjo YuEMNLP 2025 · 被引用 1 次
- Keypoint-based Progressive Chain-of-Thought Distillation for LLMsKaituo Feng, Changsheng Li, Xiaolu Zhang, Jun Zhou 等ICML 2024 · 被引用 20 次
- Improving Reasoning Capabilities in Small Models through Mixture-of-layers Distillation with Stepwise Attention on Key InformationYao Chen, Jiawei Sheng, Wenyuan Zhang, Tingwen LiuEMNLP 2025
- Teach Small Models to Reason by Curriculum DistillationWangyi Jiang, Yaojie Lu, Hongyu Lin, Xianpei Han 等EMNLP 2025
- Universal Reasoner: A Single, Composable Plug-and-Play Reasoner for Frozen LLMsJaemin Kim, Hangeol Chang, Hyunmin Hwang, Choonghan Kim 等ICML 2026 · 被引用 1 次
