Towards Efficient Pre-Trained Language Model via Feature Correlation Distillation
Kun Huang, Xin Guo, Meng Wang
摘要
Knowledge Distillation (KD) has emerged as a promising approach for compressing large Pre-trained Language Models (PLMs). The performance of KD relies on how to effectively formulate and transfer the knowledge from the teacher model to the student model. Prior arts mainly focus on directly aligning output features from the transformer block, which may impose overly strict constraints on the student model’s learning process and complicate the training process by introducing extra parameters and computational cost. Moreover, our analysis indicates that the different relations within self-attention, as adopted in other works, involves more computation complexities and can easily be constrained by the number of heads, potentially leading to suboptimal solutions. To address these issues, we propose a novel approach that builds relationships directly from output features. Specifically, we introduce token-level and sequence-level relations concurrently to fully exploit the knowledge from the teacher model. Furthermore, we propose a correlation-based distillation loss to alleviate the exact match properties inherent in traditional KL divergence or MSE loss functions. Our method, dubbed FCD, presents a simple yet effective method to compress various architectures (BERT, RoBERTa, and GPT) and model sizes (base-size and large-size). Extensive experimental results demonstrate that our distilled, smaller language models significantly surpass existing KD methods across various NLP tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- TALAS: Teacher-Anchored Layer Alignment with Adaptive Sharpness-Aware Minimization for Embedding DistillationQuoc Phong Dao, Hoang Son Nguyen, Pham Khanh Chi, Linh Ngo Van 等ACL 2026
- Skrr: Skip and Re-use Text Encoder Layers for Memory Efficient Text-to-Image GenerationHoigi Seo, Wongi Jeong, Jae-sun Seo, Se Young ChunICML 2025
它引用的顶会 Paper9
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- Similarity-Preserving Knowledge DistillationFrederick Tung, Greg MoriICCV 2019 · 被引用 1,214 次
- A Comprehensive Overhaul of Feature DistillationByeongho Heo, Jeesoo Kim, Sangdoo Yun, Hyojin Park 等ICCV 2019 · 被引用 727 次
- DynaBERT: Dynamic BERT with Adaptive Width and DepthLu Hou, Zhiqi Huang, Lifeng Shang, Xin Jiang 等NeurIPS 2020 · 被引用 401 次
- Knowledge Distillation from Internal RepresentationsGustavo Aguilar, Yuan Ling, Yu Zhang, Benjamin Z. Yao 等AAAI 2020 · 被引用 199 次
相关 Paper
- Maximizing the Effectiveness of Larger BERT Models for CompressionWen-Shu Fan, Su Lu, Shangyu Xing, Xin-Chun Li 等ACL 2025
- f-Divergence Minimization for Sequence-Level Knowledge DistillationYuqiao Wen, Zichao Li, Wenyu Du, Lili MouACL 2023 · 被引用 14 次
- Understanding and Improving Knowledge Distillation for Quantization Aware Training of Large Transformer EncodersMinsoo Kim, Sihwa Lee, Sukjin Hong, Du-Seong Chang 等EMNLP 2022 · 被引用 7 次
- Beyond Logits: Aligning Feature Dynamics for Effective Knowledge DistillationGuoqiang Gong, Jiaxing Wang, Jin Xu, Deping Xiang 等ACL 2025
- Adversarial Data Augmentation for Task-Specific Knowledge Distillation of Pre-trained TransformersMinjia Zhang, Uma-Naresh Niranjan, Yuxiong HeAAAI 2022 · 被引用 16 次
