Exploring extreme parameter compression for pre-trained language models
Benyou Wang, Yuxin Ren, Lifeng Shang, Xin Jiang, Qun Liu
Abstract
Recent work explored the potential of large-scale Transformer-based pre-trained models, especially Pre-trained Language Models (PLMs) in natural language processing. This raises many concerns from various perspectives, e.g., financial costs and carbon emissions. Compressing PLMs like BERT with negligible performance loss for faster inference and cheaper deployment has attracted much attention. In this work, we aim to explore larger compression ratios for PLMs, among which tensor decomposition is a potential but under-investigated one. Two decomposition and reconstruction protocols are further proposed to improve the effectiveness and efficiency during compression. Our compressed BERT with parameters in Transformer layers performs on-par with, sometimes slightly better than the original BERT in GLUE benchmark. A tiny version achieves performance of BERT-base with encoder parameters (i.e., less than 2M parameters excluding the embedding layer) and faster on inference. To show that the proposed method is orthogonal to existing compression methods like knowledge distillation, we also explore the benefit of the proposed method on a distilled BERT.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d6f1c308-3195-46c1-80ad-c74ea49d0301Cited by top-tier papers9
- FacT: Factor-Tuning for Lightweight Adaptation on Vision TransformerShibo Jie, Zhi-Hong DengAAAI 2023 · 182 citations
- COMCAT: Towards Efficient Compression and Customization of Attention-Based Vision ModelsJinqi Xiao, Miao Yin, Yu Gong, Xiao Zang et al.ICML 2023 · 17 citations
- Accurate Retraining-free Pruning for Pretrained Encoder-based Language ModelsSeungcheol Park, Hojun Choi, U KangICLR 2024 · 14 citations
- ReplaceMe: Network Simplification via Depth Pruning and Transformer Block LinearizationDmitriy Shopkhoev, Ammar Ali, Magauiya Zhussip, Valentin Malykh et al.NeurIPS 2025 · 10 citations
- Multi-CLS BERT: An Efficient Alternative to Traditional EnsemblingHaw-Shiuan Chang, Ruei-Yao Sun, Kathryn Ricci, Andrew McCallumACL 2023 · 8 citations
Builds on16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- Attention is not all you need: pure attention loses rank doubly exponentially with depthYihe Dong, Jean-Baptiste Cordonnier, Andreas LoukasICML 2021 · 522 citations
Related papers
- DRONE: Data-aware Low-rank Compression for Large NLP ModelsPatrick H. Chen, Hsiang-Fu Yu, Inderjit S. Dhillon, Cho-Jui HsiehNeurIPS 2021 · 109 citations
- ROSITA: Refined BERT cOmpreSsion with InTegrAted techniquesYuanxin Liu, Zheng Lin, Fengcheng YuanAAAI 2021 · 22 citations
- Differentially Private Model CompressionFatemehsadat Mireshghallah, Arturs Backurs, Huseyin A. Inan, Lukas Wutschitz et al.NeurIPS 2022 · 18 citations
- Over-parameterized Student Model via Tensor Decomposition Boosted Knowledge DistillationYu-Liang Zhan, Zhong-Yi Lu, Hao Sun, Ze-Feng GaoNeurIPS 2024 · 6 citations
- bert2BERT: Towards Reusable Pretrained Language ModelsCheng Chen, Yichun Yin, Lifeng Shang, Xin Jiang et al.ACL 2022
