Sparse Structure Search for Delta Tuning
Shengding Hu, Zhen Zhang, Ning Ding, Yadao Wang, Yasheng Wang, Zhiyuan Liu, Maosong Sun
Abstract
Adapting large pre-trained models (PTMs) through fine-tuning imposes prohibitive computational and storage burdens. Recent studies of delta tuning (DT), i.e., parameter-efficient tuning, find that only optimizing a small portion of parameters conditioned on PTMs could yield on-par performance compared to conventional fine-tuning. Generally, DT methods exquisitely design delta modules (DT modules) which could be applied to arbitrary fine-grained positions inside PTMs. However, the effectiveness of these fine-grained positions largely relies on sophisticated manual designation, thereby usually producing sub-optimal results. In contrast to the manual designation, we explore constructing DT modules in an automatic manner. We automatically S earch for the S parse S tructure of Delta Tuning (S 3 Delta). Based on a unified framework of various DT methods, S 3 Delta conducts the differentiable DT structure search through bi-level optimization and proposes shifted global sigmoid method to explicitly control the number of trainable parameters. Extensive experiments show that S 3 Delta surpasses manual and random structures with less trainable parameters. The searched structures preserve more than 99% fine-tuning performance with 0.01% trainable parameters. Moreover, the advantage of S 3 Delta is amplified with extremely low trainable parameters budgets (0.0009% ∼ 0.01%). The searched structures are transferable and explainable, providing suggestions and guidance for the future design of DT methods. Our codes are publicly available at https://github.com/thunlp/S3Delta .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- The Expressive Power of Low-Rank AdaptationYuchen Zeng, Kangwook LeeICLR 2024 · 116 citations
- G-Adapter: Towards Structure-Aware Parameter-Efficient Transfer Learning for Graph Transformer NetworksAnchun Gui, Jinqiang Ye, Han XiaoAAAI 2024 · 35 citations
- PACE: Marrying generalization in PArameter-efficient fine-tuning with Consistency rEgularizationYao Ni, Shan Zhang, Piotr KoniuszNeurIPS 2024 · 25 citations
- Random Masking Finds Winning Tickets for Parameter Efficient Fine-tuningJing Xu, Jingzhao ZhangICML 2024 · 15 citations
- PAT: Pruning-Aware Tuning for Large Language ModelsYijiang Liu, Huanrui Yang, Youxin Chen, Rongyu Zhang et al.AAAI 2025 · 1 citation
Builds on15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- On Layer Normalization in the Transformer ArchitectureRuibin Xiong, Yunchang Yang, Di He, Kai Zheng et al.ICML 2020 · 1,388 citations
- Towards a Unified View of Parameter-Efficient Transfer LearningJunxian He, Chunting Zhou, Xuezhe Ma, Taylor Berg-Kirkpatrick et al.ICLR 2022 · 1,182 citations
- Progressive Differentiable Architecture Search: Bridging the Depth Gap Between Search and EvaluationXin Chen, Lingxi Xie, Jun Wu, Qi TianICCV 2019 · 725 citations
Related papers
- DSEE: Dually Sparsity-embedded Efficient Tuning of Pre-trained Language ModelsXuxi Chen, Tianlong Chen, Weizhu Chen, Ahmed Hassan Awadallah et al.ACL 2023 · 4 citations
- Exploring the Impact of Model Scaling on Parameter-Efficient TuningYusheng Su, Chi-Min Chan, Jiali Cheng, Yujia Qin et al.EMNLP 2023 · 4 citations
- ADePT: Adaptive Decomposed Prompt Tuning for Parameter-Efficient Fine-tuningPengwei Tang, Xiaolin Hu, Yong LiuICLR 2025
- Sensitivity-Aware Efficient Fine-Tuning via Compact Dynamic-Rank AdaptationTianran Chen, Jiarui Chen, Baoquan Zhang, Zhehao Yu et al.CVPR 2025
- SPT: Learning to Selectively Insert Prompts for Better Prompt TuningWei Zhu, Ming TanEMNLP 2023 · 7 citations
