LRC-BERT: Latent-representation Contrastive Knowledge Distillation for Natural Language Understanding
Hao Fu, Shaojun Zhou, Qihong Yang, Junjie Tang, Guiquan Liu, Kaikui Liu, Xiaolong Li
摘要
The pre-training models such as BERT have achieved great results in various natural language processing problems. However, a large number of parameters need significant amounts of memory and the consumption of inference time, which makes it difficult to deploy them on edge devices. In this work, we propose a knowledge distillation method LRC-BERT based on contrastive learning to fit the output of the intermediate layer from the angular distance aspect, which is not considered by the existing distillation methods. Furthermore, we introduce a gradient perturbation-based training architecture in the training phase to increase the robustness of LRC-BERT, which is the first attempt in knowledge distillation. Additionally, in order to better capture the distribution characteristics of the intermediate layer, we design a two-stage training method for the total distillation loss. Finally, by verifying 8 datasets on the General Language Understanding Evaluation (GLUE) benchmark, the performance of the proposed LRC-BERT exceeds the existing state-of-the-art methods, which proves the effectiveness of our method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Knowledge Graph Contrastive Learning for RecommendationYuhao Yang, Chao Huang, Lianghao Xia, Chenliang LiSIGIR 2022 · 被引用 487 次
- Multi-Granularity Structural Knowledge Distillation for Language Model CompressionChang Liu, Chongyang Tao, Jiazhan Feng, Dongyan ZhaoACL 2022 · 被引用 64 次
- Finding Order in Chaos: A Novel Data Augmentation Method for Time Series in Contrastive LearningBerken Utku Demirel, Christian HolzNeurIPS 2023 · 被引用 48 次
- Learning to Distill Global Representation for Sparse-View CTZilong Li, Chenglong Ma, Jie Chen, Junping Zhang 等ICCV 2023 · 被引用 22 次
- A Good Learner can Teach Better: Teacher-Student Collaborative Knowledge DistillationAyan Sengupta, Shantanu Dixit, Md. Shad Akhtar, Tanmoy ChakrabortyICLR 2024 · 被引用 16 次
它引用的顶会 Paper6
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersWenhui Wang, Furu Wei, Li Dong, Hangbo Bao 等NeurIPS 2020 · 被引用 2,727 次
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 被引用 1,305 次
- MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited DevicesZhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu 等ACL 2020 · 被引用 660 次
- Knowledge Distillation from Internal RepresentationsGustavo Aguilar, Yuan Ling, Yu Zhang, Benjamin Z. Yao 等AAAI 2020 · 被引用 199 次
相关 Paper
- Contrastive Distillation on Intermediate Representations for Language Model CompressionSiqi Sun, Zhe Gan, Yuwei Fang, Yu Cheng 等EMNLP 2020 · 被引用 59 次
- BERT-EMD: Many-to-Many Layer Mapping for BERT Compression with Earth Mover's DistanceJianquan Li, Xiaokang Liu, Honghong Zhao, Ruifeng Xu 等EMNLP 2020 · 被引用 44 次
- Universal-KD: Attention-based Output-Grounded Intermediate Layer Knowledge DistillationYimeng Wu, Mehdi Rezagholizadeh, Abbas Ghaddar, Md. Akmal Haidar 等EMNLP 2021 · 被引用 18 次
- TernaryBERT: Distillation-aware Ultra-low Bit BERTWei Zhang, Lu Hou, Yichun Yin, Lifeng Shang 等EMNLP 2020 · 被引用 147 次
- SKDBERT: Compressing BERT via Stochastic Knowledge DistillationZixiang Ding, Guoqing Jiang, Shuai Zhang, Lin Guo 等AAAI 2023 · 被引用 13 次
