Enabling Lightweight Fine-tuning for Pre-trained Language Model Compression based on Matrix Product Operators
Peiyu Liu, Ze-Feng Gao, Wayne Xin Zhao, Zhi-Yuan Xie, Zhong-Yi Lu, Ji-Rong Wen
摘要
This paper presents a novel pre-trained language models (PLM) compression approach based on the matrix product operator (short as MPO) from quantum many-body physics. It can decompose an original matrix into central tensors (containing the core information) and auxiliary tensors (with only a small proportion of parameters). With the decomposed MPO structure, we propose a novel fine-tuning strategy by only updating the parameters from the auxiliary tensors, and design an optimization algorithm for MPO-based approximation over stacked network architectures. Our approach can be applied to the original or the compressed PLMs in a general way, which derives a lighter network and significantly reduces the parameters to be fine-tuned. Extensive experiments have demonstrated the effectiveness of the proposed approach in model compression, especially the reduction in finetuning parameters (91% reduction on average). The code to reproduce the results of this paper can be found at https://github.com/ RUCAIBox/MPOP .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Sparse Low-rank Adaptation of Pre-trained Language ModelsNing Ding, Xingtai Lv, Qiaosen Wang, Yulin Chen 等EMNLP 2023 · 被引用 42 次
- QuanTA: Efficient High-Rank Fine-Tuning of LLMs with Quantum-Informed Tensor AdaptationZhuo Chen, Rumen Dangovski, Charlotte Loh, Owen Dugan 等NeurIPS 2024 · 被引用 38 次
- Exploring extreme parameter compression for pre-trained language modelsBenyou Wang, Yuxin Ren, Lifeng Shang, Xin Jiang 等ICLR 2022 · 被引用 23 次
- On the Impact of Knowledge Distillation for Model InterpretabilityHyeongrok Han, Siwon Kim, Hyun-Soo Choi, Sungroh YoonICML 2023 · 被引用 13 次
- LaX: Boosting Low-Rank Training of Foundation Models via Latent CrossingRuijie Zhang, Ziyue Liu, Zhengyang Wang, Zheng ZhangNeurIPS 2025 · 被引用 7 次
它引用的顶会 Paper9
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited DevicesZhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu 等ACL 2020 · 被引用 660 次
- DynaBERT: Dynamic BERT with Adaptive Width and DepthLu Hou, Zhiqi Huang, Lifeng Shang, Xin Jiang 等NeurIPS 2020 · 被引用 401 次
- FastBERT: a Self-distilling BERT with Adaptive Inference TimeWeijie Liu, Peng Zhou, Zhiruo Wang, Zhe Zhao 等ACL 2020 · 被引用 257 次
相关 Paper
- Small Pre-trained Language Models Can be Fine-tuned as Large Models via Over-ParameterizationZe-Feng Gao, Kun Zhou, Peiyu Liu, Wayne Xin Zhao 等ACL 2023 · 被引用 3 次
- NTK-approximating MLP Fusion for Efficient Language Model Fine-tuningTianxin Wei, Zeming Guo, Yifan Chen, Jingrui HeICML 2023
- Pruning Pre-trained Language Models Without Fine-TuningTing Jiang, Deqing Wang, Fuzhen Zhuang, Ruobing Xie 等ACL 2023 · 被引用 1 次
- Towards Robust Low-Resource Fine-Tuning with Multi-View Compressed RepresentationsLinlin Liu, Xingxuan Li, Megh Thakkar, Xin Li 等ACL 2023 · 被引用 3 次
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersWenhui Wang, Furu Wei, Li Dong, Hangbo Bao 等NeurIPS 2020 · 被引用 2,727 次
