AD-KD: Attribution-Driven Knowledge Distillation for Language Model Compression
Siyue Wu, Hongzhan Chen, Xiaojun Quan, Qifan Wang, Rui Wang
摘要
Knowledge distillation has attracted a great deal of interest recently to compress pre-trained language models. However, existing knowledge distillation methods suffer from two limitations. First, the student model simply imitates the teacher's behavior while ignoring the underlying reasoning. Second, these methods usually focus on the transfer of sophisticated model-specific knowledge but overlook dataspecific knowledge. In this paper, we present a novel attribution-driven knowledge distillation approach, which explores the token-level rationale behind the teacher model based on Integrated Gradients (IG) and transfers attribution knowledge to the student model. To enhance the knowledge transfer of model reasoning and generalization, we further explore multi-view attribution distillation on all potential decisions of the teacher. Comprehensive experiments are conducted with BERT on the GLUE benchmark. The experimental results demonstrate the superior performance of our approach to several state-of-the-art methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Over-parameterized Student Model via Tensor Decomposition Boosted Knowledge DistillationYu-Liang Zhan, Zhong-Yi Lu, Hao Sun, Ze-Feng GaoNeurIPS 2024 · 被引用 6 次
- FuseChat: Knowledge Fusion of Chat ModelsFanqi Wan, Longguang Zhong, Ziyi Yang, Ruijun Chen 等EMNLP 2025 · 被引用 4 次
- GrAInS: Gradient-based Attribution for Inference-Time Steering of LLMs and VLMsDuy Nguyen, Archiki Prasad, Elias Stengel-Eskin, Mohit BansalACL 2026 · 被引用 4 次
- DGS-Net: Distillation-Guided Gradient Surgery for CLIP Fine-Tuning in AI-Generated Image DetectionJiazhen Yan, Ziqiang Li, Fan Wang, Boyu Wang 等ICML 2026 · 被引用 1 次
- Late-to-Early Training: LET LLMs Learn Earlier, So Faster and BetterJi Zhao, Shitong Shao, Yufei Gu, Xun Zhou 等ICLR 2026 · 被引用 1 次
它引用的顶会 Paper13
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersWenhui Wang, Furu Wei, Li Dong, Hangbo Bao 等NeurIPS 2020 · 被引用 2,727 次
- MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited DevicesZhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu 等ACL 2020 · 被引用 660 次
- Restricting the Flow: Information Bottlenecks for AttributionKarl Schulz, Leon Sixt, Federico Tombari, Tim LandgrafICLR 2020 · 被引用 220 次
- On Identifiability in TransformersGino Brunner, Yang Liu, Damian Pascual, Oliver Richter 等ICLR 2020 · 被引用 210 次
相关 Paper
- Explain in Your Own Words: Improving Reasoning via Token-Selective Dual Knowledge DistillationMinsang Kim, Seung Jun BaekICLR 2026 · 被引用 15 次
- Multi-Granularity Structural Knowledge Distillation for Language Model CompressionChang Liu, Chongyang Tao, Jiazhan Feng, Dongyan ZhaoACL 2022 · 被引用 64 次
- Towards Efficient Pre-Trained Language Model via Feature Correlation DistillationKun Huang, Xin Guo, Meng WangNeurIPS 2023 · 被引用 8 次
- Maximizing the Effectiveness of Larger BERT Models for CompressionWen-Shu Fan, Su Lu, Shangyu Xing, Xin-Chun Li 等ACL 2025
- How to Trade Off the Quantity and Capacity of Teacher Ensemble: Learning Categorical Distribution to Stochastically Employ a Teacher for DistillationZixiang Ding, Guoqing Jiang, Shuai Zhang, Lin Guo 等AAAI 2024 · 被引用 4 次
