Sparse Structure Search for Delta Tuning
Shengding Hu, Zhen Zhang, Ning Ding, Yadao Wang, Yasheng Wang, Zhiyuan Liu, Maosong Sun
摘要
Adapting large pre-trained models (PTMs) through fine-tuning imposes prohibitive computational and storage burdens. Recent studies of delta tuning (DT), i.e., parameter-efficient tuning, find that only optimizing a small portion of parameters conditioned on PTMs could yield on-par performance compared to conventional fine-tuning. Generally, DT methods exquisitely design delta modules (DT modules) which could be applied to arbitrary fine-grained positions inside PTMs. However, the effectiveness of these fine-grained positions largely relies on sophisticated manual designation, thereby usually producing sub-optimal results. In contrast to the manual designation, we explore constructing DT modules in an automatic manner. We automatically S earch for the S parse S tructure of Delta Tuning (S 3 Delta). Based on a unified framework of various DT methods, S 3 Delta conducts the differentiable DT structure search through bi-level optimization and proposes shifted global sigmoid method to explicitly control the number of trainable parameters. Extensive experiments show that S 3 Delta surpasses manual and random structures with less trainable parameters. The searched structures preserve more than 99% fine-tuning performance with 0.01% trainable parameters. Moreover, the advantage of S 3 Delta is amplified with extremely low trainable parameters budgets (0.0009% ∼ 0.01%). The searched structures are transferable and explainable, providing suggestions and guidance for the future design of DT methods. Our codes are publicly available at https://github.com/thunlp/S3Delta .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- The Expressive Power of Low-Rank AdaptationYuchen Zeng, Kangwook LeeICLR 2024 · 被引用 116 次
- G-Adapter: Towards Structure-Aware Parameter-Efficient Transfer Learning for Graph Transformer NetworksAnchun Gui, Jinqiang Ye, Han XiaoAAAI 2024 · 被引用 35 次
- PACE: Marrying generalization in PArameter-efficient fine-tuning with Consistency rEgularizationYao Ni, Shan Zhang, Piotr KoniuszNeurIPS 2024 · 被引用 25 次
- Random Masking Finds Winning Tickets for Parameter Efficient Fine-tuningJing Xu, Jingzhao ZhangICML 2024 · 被引用 15 次
- PAT: Pruning-Aware Tuning for Large Language ModelsYijiang Liu, Huanrui Yang, Youxin Chen, Rongyu Zhang 等AAAI 2025 · 被引用 1 次
它引用的顶会 Paper15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- On Layer Normalization in the Transformer ArchitectureRuibin Xiong, Yunchang Yang, Di He, Kai Zheng 等ICML 2020 · 被引用 1,388 次
- Towards a Unified View of Parameter-Efficient Transfer LearningJunxian He, Chunting Zhou, Xuezhe Ma, Taylor Berg-Kirkpatrick 等ICLR 2022 · 被引用 1,182 次
- Progressive Differentiable Architecture Search: Bridging the Depth Gap Between Search and EvaluationXin Chen, Lingxi Xie, Jun Wu, Qi TianICCV 2019 · 被引用 725 次
相关 Paper
- DSEE: Dually Sparsity-embedded Efficient Tuning of Pre-trained Language ModelsXuxi Chen, Tianlong Chen, Weizhu Chen, Ahmed Hassan Awadallah 等ACL 2023 · 被引用 4 次
- Exploring the Impact of Model Scaling on Parameter-Efficient TuningYusheng Su, Chi-Min Chan, Jiali Cheng, Yujia Qin 等EMNLP 2023 · 被引用 4 次
- ADePT: Adaptive Decomposed Prompt Tuning for Parameter-Efficient Fine-tuningPengwei Tang, Xiaolin Hu, Yong LiuICLR 2025
- Sensitivity-Aware Efficient Fine-Tuning via Compact Dynamic-Rank AdaptationTianran Chen, Jiarui Chen, Baoquan Zhang, Zhehao Yu 等CVPR 2025
- SPT: Learning to Selectively Insert Prompts for Better Prompt TuningWei Zhu, Ming TanEMNLP 2023 · 被引用 7 次
