DRONE: Data-aware Low-rank Compression for Large NLP Models
Patrick H. Chen, Hsiang-Fu Yu, Inderjit S. Dhillon, Cho-Jui Hsieh
摘要
The representations learned by large-scale NLP models such as BERT have been widely used in various tasks. However, the increasing model size of the pre-trained models also brings efficiency challenges, including inference speed and model size when deploying models on mobile devices. Specifically, most operations in BERT consist of matrix multiplications. These matrices are not low-rank and thus canonical matrix decompositions do not lead to efficient approximations. In this paper, we observe that the learned representation of each layer lies in a lowdimensional space. Based on this observation, we propose DRONE (data-aware low-rank compression), a provably optimal low-rank decomposition of weight matrices, which has a simple closed form solution that can be efficiently computed. DRONE can be applied to both fully-connected and self-attention layers appearing in the BERT model. In addition to compressing standard models, our method can also be used on distilled BERT models to further improve the compression rate. Experimental results show that DRONE is able to improve both model size and inference speed with limited loss in accuracy. Specifically, DRONE alone achieves 1.92x speedup on the MRPC task with only 1.5% loss in accuracy, and when DRONE is combined with distillation, it further achieves over 12.3x speedup on various natural language inference tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper23
- A Survey on Model Compression and Acceleration for Pretrained Language ModelsCanwen Xu, Julian J. McAuleyAAAI 2023 · 被引用 96 次
- LQ-LoRA: Low-rank plus Quantized Matrix Decomposition for Efficient Language Model FinetuningHan Guo, Philip Greengard, Eric P. Xing, Yoon KimICLR 2024 · 被引用 94 次
- Rank Diminishing in Deep Neural NetworksRuili Feng, Kecheng Zheng, Yukun Huang, Deli Zhao 等NeurIPS 2022 · 被引用 64 次
- Hypernetwork-based Meta-Learning for Low-Rank Physics-Informed Neural NetworksWoojin Cho, Kookjin Lee, Donsub Rim, Noseong ParkNeurIPS 2023 · 被引用 62 次
- Compressing Transformers: Features Are Low-Rank, but Weights Are Not!Hao Yu, Jianxin WuAAAI 2023 · 被引用 59 次
它引用的顶会 Paper9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited DevicesZhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu 等ACL 2020 · 被引用 660 次
- Movement Pruning: Adaptive Sparsity by Fine-TuningVictor Sanh, Thomas Wolf, Alexander M. RushNeurIPS 2020 · 被引用 656 次
- HAWQ: Hessian AWare Quantization of Neural Networks With Mixed-PrecisionZhen Dong, Zhewei Yao, Amir Gholami, Michael W. Mahoney 等ICCV 2019 · 被引用 645 次
- The Lottery Ticket Hypothesis for Pre-trained BERT NetworksTianlong Chen, Jonathan Frankle, Shiyu Chang, Sijia Liu 等NeurIPS 2020 · 被引用 428 次
相关 Paper
- Exploring extreme parameter compression for pre-trained language modelsBenyou Wang, Yuxin Ren, Lifeng Shang, Xin Jiang 等ICLR 2022 · 被引用 23 次
- NAS-BERT: Task-Agnostic and Adaptive-Size BERT Compression with Neural Architecture SearchJin Xu, Xu Tan, Renqian Luo, Kaitao Song 等KDD 2021 · 被引用 49 次
- XtremeDistil: Multi-stage Distillation for Massive Multilingual ModelsSubhabrata Mukherjee, Ahmed Hassan AwadallahACL 2020 · 被引用 4 次
- The Optimal BERT Surgeon: Scalable and Accurate Second-Order Pruning for Large Language ModelsEldar Kurtic, Daniel Campos, Tuan Nguyen, Elias Frantar 等EMNLP 2022 · 被引用 4 次
- BERT-EMD: Many-to-Many Layer Mapping for BERT Compression with Earth Mover's DistanceJianquan Li, Xiaokang Liu, Honghong Zhao, Ruifeng Xu 等EMNLP 2020 · 被引用 44 次
