Fine-Tuning Pre-Trained Language Models Effectively by Optimizing Subnetworks Adaptively
Haojie Zhang, Ge Li, Jia Li, Zhongjin Zhang, Yuqi Zhu, Zhi Jin
Abstract
Large-scale pre-trained language models have achieved impressive results on a wide range of downstream tasks recently. However, fine-tuning an extremely large-scale pre-trained language model on limited target datasets is often plagued by overfitting and representation degradation. In this paper, we propose a Dynamic Parameter Selection (DPS) algorithm for the large-scale pre-trained models during fine-tuning, which adaptively selects a more promising subnetwork to perform staging updates based on gradients of back-propagation. Experiments on the GLUE benchmark show that DPS outperforms previous fine-tuning methods in terms of overall performance and stability, and consistently achieves better results with variable pre-trained language models. In addition, DPS brings a large magnitude of improvement in out-of-domain transferring experiments and low-resource scenarios, which shows that it can maintain stable general contextual features and reduce the representation collapse. We release our code at https://github.com/ZhangHaojie077/DPS .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Fine-Tuned Large Language Model for Visualization System: A Study on Self-Regulated Learning in EducationLin Gao, Jing Lu, Zekai Shao, Ziyue Lin et al.IEEE VIS 2024 · 27 citations
- LiNeS: Post-training Layer Scaling Prevents Forgetting and Enhances Model MergingKe Wang, Nikolaos Dimitriadis, Alessandro Favero, Guillermo Ortiz-Jiménez et al.ICLR 2025
- MODEL SOUPS NEED ONLY ONE INGREDIENTAlireza Abdollahpourrostam, Nikolaos Dimitriadis, Adam Hazimeh, Pascal FrossardICML 2026
Builds on11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- R-Drop: Regularized Dropout for Neural NetworksXiaobo Liang, Lijun Wu, Juntao Li, Yue Wang et al.NeurIPS 2021 · 610 citations
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
- On the Stability of Fine-tuning BERT: Misconceptions, Explanations, and Strong BaselinesMarius Mosbach, Maksym Andriushchenko, Dietrich KlakowICLR 2021 · 448 citations
Related papers
- Raise a Child in Large Language Model: Towards Effective and Generalizable Fine-tuningRunxin Xu, Fuli Luo, Zhiyuan Zhang, Chuanqi Tan et al.EMNLP 2021 · 129 citations
- ROSE: Robust Selective Fine-tuning for Pre-trained Language ModelsLan Jiang, Hao Zhou, Yankai Lin, Peng Li et al.EMNLP 2022 · 5 citations
- HyPe: Better Pre-trained Language Model Fine-tuning with Hidden Representation PerturbationHongyi Yuan, Zheng Yuan, Chuanqi Tan, Fei Huang et al.ACL 2023 · 5 citations
- An Empirical Study on Hyperparameter Optimization for Fine-Tuning Pre-trained Language ModelsXueqing Liu, Chi WangACL 2021
- HiFi: High-Information Attention Heads Hold for Parameter-Efficient Model AdaptationAnchun Gui, Han XiaoACL 2023 · 1 citation
