One Network, Many Masks: Towards More Parameter-Efficient Transfer Learning
Guangtao Zeng, Peiyuan Zhang, Wei Lu
Abstract
Fine-tuning pre-trained language models for multiple tasks tends to be expensive in terms of storage. To mitigate this, parameter-efficient transfer learning (PETL) methods have been proposed to address this issue, but they still require a significant number of parameters and storage when being applied to broader ranges of tasks. To achieve even greater storage reduction, we propose PROPETL, a novel method that enables efficient sharing of a single PETL module which we call prototype network (e.g., adapter, LoRA, and prefix-tuning) across layers and tasks. We then learn binary masks to select different sub-networks from the shared prototype network and apply them as PETL modules into different layers. We find that the binary masks can determine crucial information from the network, which is often ignored in previous studies. Our work can also be seen as a type of pruning method, where we find that overparameterization also exists in the seemingly small PETL modules. We evaluate PROPETL on various downstream tasks and show that it can outperform other PETL methods with approximately 10% of the parameter storage required by the latter. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- DoRA: Enhancing Parameter-Efficient Fine-Tuning with Dynamic Rank DistributionYulong Mao, Kaiyu Huang, Changhao Guan, Ganglin Bao et al.ACL 2024 · 15 citations
- Increasing Model Capacity for Free: A Simple Strategy for Parameter Efficient Fine-tuningHaobo Song, Hao Zhao, Soumajit Majumder, Tao LinICLR 2024 · 11 citations
- Prototype-based HyperAdapter for Sample-Efficient Multi-task TuningHao Zhao, Jie Fu, Zhaofeng HeEMNLP 2023 · 3 citations
- SliceFine: The Universal Winning-Slice Hypothesis for Pretrained NetworksMd Kowsher, Ali Polat, Ehsan Ardehaly, Mehrdad Salehi et al.ICML 2026
- RoCoFT: Efficient Finetuning of Large Language Models with Row-Column UpdatesMd. Kowsher, Tara Esmaeilbeig, Chun-Nam Yu, Chen Chen et al.ACL 2025
Builds on10
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- UniPELT: A Unified Framework for Parameter-Efficient Language Model TuningYuning Mao, Lambert Mathias, Rui Hou, Amjad Almahairi et al.ACL 2022 · 225 citations
- Revisiting Few-sample BERT Fine-tuningTianyi Zhang, Felix Wu, Arzoo Katiyar, Kilian Q. Weinberger et al.ICLR 2021 · 172 citations
- Raise a Child in Large Language Model: Towards Effective and Generalizable Fine-tuningRunxin Xu, Fuli Luo, Zhiyuan Zhang, Chuanqi Tan et al.EMNLP 2021 · 129 citations
Related papers
- LST: Ladder Side-Tuning for Parameter and Memory Efficient Transfer LearningYi-Lin Sung, Jaemin Cho, Mohit BansalNeurIPS 2022 · 347 citations
- From Bottom to Top: Extending the Potential of Parameter Efficient Fine-TuningJihao Gu, Zelin Wang, Yibo Zhang, Ziji Zhang et al.EMNLP 2024 · 3 citations
- Faster Parameter-Efficient Tuning with Token Redundancy ReductionKwonyoung Kim, Jungin Park, Jin Kim, Hyeongjun Kwon et al.CVPR 2025
- MTLoRA: A Low-Rank Adaptation Approach for Efficient Multi-Task LearningAhmed Agiza, Marina Neseem, Sherief RedaCVPR 2024
- UniPT: Universal Parallel Tuning for Transfer Learning with Efficient Parameter and MemoryHaiwen Diao, Bo Wan, Ying Zhang, Xu Jia et al.CVPR 2024
