Multi-Granularity Structural Knowledge Distillation for Language Model Compression
Chang Liu, Chongyang Tao, Jiazhan Feng, Dongyan Zhao
摘要
Transferring the knowledge to a small model through distillation has raised great interest in recent years. Prevailing methods transfer the knowledge derived from mono-granularity language units (e.g., token-level or sample-level), which is not enough to represent the rich semantics of a text and may lose some vital knowledge. Besides, these methods form the knowledge as individual representations or their simple dependencies, neglecting abundant structural relations among intermediate representations. To overcome the problems, we present a novel knowledge distillation framework that gathers intermediate representations from multiple semantic granularities (e.g., tokens, spans and samples) and forms the knowledge as more sophisticated structural relations specified as the pair-wise interactions and the triplet-wise geometric angles based on multi-granularity representations. Moreover, we propose distilling the well-organized multi-granularity structural knowledge to the student hierarchically across layers. Experimental results on GLUE benchmark demonstrate that our method outperforms advanced distillation methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- A Survey on Model Compression and Acceleration for Pretrained Language ModelsCanwen Xu, Julian J. McAuleyAAAI 2023 · 被引用 96 次
- Fine-Grained Distillation for Long Document RetrievalYucheng Zhou, Tao Shen, Xiubo Geng, Chongyang Tao 等AAAI 2024 · 被引用 44 次
- Multi-Level Optimal Transport for Universal Cross-Tokenizer Knowledge Distillation on Language ModelsXiao Cui, Mo Zhu, Yulei Qin, Liang Xie 等AAAI 2025 · 被引用 31 次
- SelecTKD: Selective Token-Weighted Knowledge Distillation for LLMsHaiduo Huang, Jiangcheng Song, Yadong Zhang, Pengju RenCVPR 2026 · 被引用 18 次
- LaKD: Length-agnostic Knowledge Distillation for Trajectory Prediction with Any Length ObservationsYuhang Li, Changsheng Li, Ruilin Lv, Rongqing Li 等NeurIPS 2024 · 被引用 16 次
它引用的顶会 Paper18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersWenhui Wang, Furu Wei, Li Dong, Hangbo Bao 等NeurIPS 2020 · 被引用 2,727 次
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 被引用 1,305 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
相关 Paper
- MTA: Multi-Granular Trajectory Alignment for Large Language Model DistillationPham Khanh Chi, Quoc Phong Dao, Thuat Nguyen, Linh Ngo Van 等ACL 2026
- Complementary Relation Contrastive DistillationJinguo Zhu, Shixiang Tang, Dapeng Chen, Shijie Yu 等CVPR 2021
- AD-KD: Attribution-Driven Knowledge Distillation for Language Model CompressionSiyue Wu, Hongzhan Chen, Xiaojun Quan, Qifan Wang 等ACL 2023 · 被引用 10 次
- Multi-Label Knowledge DistillationPenghui Yang, Ming-Kun Xie, Chen-Chen Zong, Lei Feng 等ICCV 2023 · 被引用 16 次
- Knowledge Distillation from Internal RepresentationsGustavo Aguilar, Yuan Ling, Yu Zhang, Benjamin Z. Yao 等AAAI 2020 · 被引用 199 次
