Distilling Linguistic Context for Language Model Compression
Geondo Park, Gyeongman Kim, Eunho Yang
摘要
A computationally expensive and memory intensive neural network lies behind the recent success of language representation learning. Knowledge distillation, a major technique for deploying such a vast language model in resource-scarce environments, transfers the knowledge on individual word representations learned without restrictions. In this paper, inspired by the recent observations that language representations are relatively positioned and have more semantic knowledge as a whole, we present a new knowledge distillation objective for language representation learning that transfers the contextual knowledge via two types of relationships across representations: Word Relation and Layer Transforming Relation. Unlike other recent distillation techniques for the language models, our contextual distillation does not have any restrictions on architectural changes between teacher and student. We validate the effectiveness of our method on challenging benchmarks of language understanding tasks, not only in architectures of various sizes, but also in combination with DynaBERT, the recently proposed adaptive size pruning method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Multi-Granularity Structural Knowledge Distillation for Language Model CompressionChang Liu, Chongyang Tao, Jiazhan Feng, Dongyan ZhaoACL 2022 · 被引用 64 次
- Egeria: Efficient DNN Training with Knowledge-Guided Layer FreezingYiding Wang, Decang Sun, Kai Chen, Fan Lai 等EuroSys 2023 · 被引用 43 次
- Multi-Level Optimal Transport for Universal Cross-Tokenizer Knowledge Distillation on Language ModelsXiao Cui, Mo Zhu, Yulei Qin, Liang Xie 等AAAI 2025 · 被引用 31 次
- A Good Learner can Teach Better: Teacher-Student Collaborative Knowledge DistillationAyan Sengupta, Shantanu Dixit, Md. Shad Akhtar, Tanmoy ChakrabortyICLR 2024 · 被引用 16 次
- Fast and Robust Early-Exiting Framework for Autoregressive Language Models with Synchronized Parallel DecodingSangmin Bae, Jongwoo Ko, Hwanjun Song, Se-Young YunEMNLP 2023 · 被引用 12 次
它引用的顶会 Paper6
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersWenhui Wang, Furu Wei, Li Dong, Hangbo Bao 等NeurIPS 2020 · 被引用 2,727 次
- MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited DevicesZhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu 等ACL 2020 · 被引用 660 次
- DynaBERT: Dynamic BERT with Adaptive Width and DepthLu Hou, Zhiqi Huang, Lifeng Shang, Xin Jiang 等NeurIPS 2020 · 被引用 401 次
- FastBERT: a Self-distilling BERT with Adaptive Inference TimeWeijie Liu, Peng Zhou, Zhiruo Wang, Zhe Zhao 等ACL 2020 · 被引用 257 次
- On Identifiability in TransformersGino Brunner, Yang Liu, Damian Pascual, Oliver Richter 等ICLR 2020 · 被引用 210 次
相关 Paper
- MTA: Multi-Granular Trajectory Alignment for Large Language Model DistillationPham Khanh Chi, Quoc Phong Dao, Thuat Nguyen, Linh Ngo Van 等ACL 2026
- Towards Efficient Pre-Trained Language Model via Feature Correlation DistillationKun Huang, Xin Guo, Meng WangNeurIPS 2023 · 被引用 8 次
- Multi-level Distillation of Semantic Knowledge for Pre-training Multilingual Language ModelMingqi Li, Fei Ding, Dan Zhang, Long Cheng 等EMNLP 2022 · 被引用 3 次
- SKDBERT: Compressing BERT via Stochastic Knowledge DistillationZixiang Ding, Guoqing Jiang, Shuai Zhang, Lin Guo 等AAAI 2023 · 被引用 13 次
- How to Trade Off the Quantity and Capacity of Teacher Ensemble: Learning Categorical Distribution to Stochastically Employ a Teacher for DistillationZixiang Ding, Guoqing Jiang, Shuai Zhang, Lin Guo 等AAAI 2024 · 被引用 4 次
