A Win-win Deal: Towards Sparse and Robust Pre-trained Language Models
Yuanxin Liu, Fandong Meng, Zheng Lin, Jiangnan Li, Peng Fu, Yanan Cao, Weiping Wang, Jie Zhou
Abstract
Despite the remarkable success of pre-trained language models (PLMs), they still face two challenges: First, large-scale PLMs are inefficient in terms of memory footprint and computation. Second, on the downstream tasks, PLMs tend to rely on the dataset bias and struggle to generalize to out-of-distribution (OOD) data. In response to the efficiency problem, recent studies show that dense PLMs can be replaced with sparse subnetworks without hurting the performance. Such subnetworks can be found in three scenarios: 1) the fine-tuned PLMs, 2) the raw PLMs and then fine-tuned in isolation, and even inside 3) PLMs without any parameter fine-tuning. However, these results are only obtained in the in-distribution (ID) setting. In this paper, we extend the study on PLMs subnetworks to the OOD setting, investigating whether sparsity and robustness to dataset bias can be achieved simultaneously. To this end, we conduct extensive experiments with the pre-trained BERT model on three natural language understanding (NLU) tasks. Our results demonstrate that sparse and robust subnetworks (SRNets) can consistently be found in BERT, across the aforementioned three scenarios, using different training and compression methods. Furthermore, we explore the upper bound of SRNets using the OOD information and show that there exist sparse and almost unbiased BERT subnetworks. Finally, we present 1) an analytical study that provides insights on how to promote the efficiency of SRNets searching process and 2) a solution to improve subnetworks' performance at high sparsity. The code is available at https://github.com/llyx97/sparse-and-robust-PLM.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Task-Specific Skill Localization in Fine-tuned Language ModelsAbhishek Panigrahi, Nikunj Saunshi, Haoyu Zhao, Sanjeev AroraICML 2023 · 100 citations
- UPop: Unified and Progressive Pruning for Compressing Vision-Language TransformersDachuan Shi, Chaofan Tao, Ying Jin, Zhendong Yang et al.ICML 2023 · 64 citations
- Reinforcement Learning Finetunes Small Subnetworks in Large Language ModelsSagnik Mukherjee, Lifan Yuan, Dilek Hakkani-Tur, Hao PengNeurIPS 2025 · 43 citations
Builds on18
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Movement Pruning: Adaptive Sparsity by Fine-TuningVictor Sanh, Thomas Wolf, Alexander M. RushNeurIPS 2020 · 656 citations
- The Lottery Ticket Hypothesis for Pre-trained BERT NetworksTianlong Chen, Jonathan Frankle, Shiyu Chang, Sijia Liu et al.NeurIPS 2020 · 428 citations
- HYDRA: Pruning Adversarially Robust Neural NetworksVikash Sehwag, Shiqi Wang, Prateek Mittal, Suman JanaNeurIPS 2020 · 242 citations
- TernaryBERT: Distillation-aware Ultra-low Bit BERTWei Zhang, Lu Hou, Yichun Yin, Lifeng Shang et al.EMNLP 2020 · 147 citations
Related papers
- Compressing and Debiasing Vision-Language Pre-Trained Models for Visual Question AnsweringQingyi Si, Yuanxin Liu, Zheng Lin, Peng Fu et al.EMNLP 2023 · 2 citations
- NAS-BERT: Task-Agnostic and Adaptive-Size BERT Compression with Neural Architecture SearchJin Xu, Xu Tan, Renqian Luo, Kaitao Song et al.KDD 2021 · 49 citations
- DSEE: Dually Sparsity-embedded Efficient Tuning of Pre-trained Language ModelsXuxi Chen, Tianlong Chen, Weizhu Chen, Ahmed Hassan Awadallah et al.ACL 2023 · 4 citations
- DynaBERT: Dynamic BERT with Adaptive Width and DepthLu Hou, Zhiqi Huang, Lifeng Shang, Xin Jiang et al.NeurIPS 2020 · 401 citations
- IsoBN: Fine-Tuning BERT with Isotropic Batch NormalizationWenxuan Zhou, Bill Yuchen Lin, Xiang RenAAAI 2021 · 29 citations
