LRC-BERT: Latent-representation Contrastive Knowledge Distillation for Natural Language Understanding
Hao Fu, Shaojun Zhou, Qihong Yang, Junjie Tang, Guiquan Liu, Kaikui Liu, Xiaolong Li
Abstract
The pre-training models such as BERT have achieved great results in various natural language processing problems. However, a large number of parameters need significant amounts of memory and the consumption of inference time, which makes it difficult to deploy them on edge devices. In this work, we propose a knowledge distillation method LRC-BERT based on contrastive learning to fit the output of the intermediate layer from the angular distance aspect, which is not considered by the existing distillation methods. Furthermore, we introduce a gradient perturbation-based training architecture in the training phase to increase the robustness of LRC-BERT, which is the first attempt in knowledge distillation. Additionally, in order to better capture the distribution characteristics of the intermediate layer, we design a two-stage training method for the total distillation loss. Finally, by verifying 8 datasets on the General Language Understanding Evaluation (GLUE) benchmark, the performance of the proposed LRC-BERT exceeds the existing state-of-the-art methods, which proves the effectiveness of our method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers13
- Knowledge Graph Contrastive Learning for RecommendationYuhao Yang, Chao Huang, Lianghao Xia, Chenliang LiSIGIR 2022 · 487 citations
- Multi-Granularity Structural Knowledge Distillation for Language Model CompressionChang Liu, Chongyang Tao, Jiazhan Feng, Dongyan ZhaoACL 2022 · 64 citations
- Finding Order in Chaos: A Novel Data Augmentation Method for Time Series in Contrastive LearningBerken Utku Demirel, Christian HolzNeurIPS 2023 · 48 citations
- Learning to Distill Global Representation for Sparse-View CTZilong Li, Chenglong Ma, Jie Chen, Junping Zhang et al.ICCV 2023 · 22 citations
- A Good Learner can Teach Better: Teacher-Student Collaborative Knowledge DistillationAyan Sengupta, Shantanu Dixit, Md. Shad Akhtar, Tanmoy ChakrabortyICLR 2024 · 16 citations
Builds on6
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersWenhui Wang, Furu Wei, Li Dong, Hangbo Bao et al.NeurIPS 2020 · 2,727 citations
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 1,305 citations
- MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited DevicesZhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu et al.ACL 2020 · 660 citations
- Knowledge Distillation from Internal RepresentationsGustavo Aguilar, Yuan Ling, Yu Zhang, Benjamin Z. Yao et al.AAAI 2020 · 199 citations
Related papers
- Contrastive Distillation on Intermediate Representations for Language Model CompressionSiqi Sun, Zhe Gan, Yuwei Fang, Yu Cheng et al.EMNLP 2020 · 59 citations
- BERT-EMD: Many-to-Many Layer Mapping for BERT Compression with Earth Mover's DistanceJianquan Li, Xiaokang Liu, Honghong Zhao, Ruifeng Xu et al.EMNLP 2020 · 44 citations
- Universal-KD: Attention-based Output-Grounded Intermediate Layer Knowledge DistillationYimeng Wu, Mehdi Rezagholizadeh, Abbas Ghaddar, Md. Akmal Haidar et al.EMNLP 2021 · 18 citations
- TernaryBERT: Distillation-aware Ultra-low Bit BERTWei Zhang, Lu Hou, Yichun Yin, Lifeng Shang et al.EMNLP 2020 · 147 citations
- SKDBERT: Compressing BERT via Stochastic Knowledge DistillationZixiang Ding, Guoqing Jiang, Shuai Zhang, Lin Guo et al.AAAI 2023 · 13 citations
