Learn-to-Share: A Hardware-friendly Transfer Learning Framework Exploiting Computation and Parameter Sharing
Cheng Fu, Hanxian Huang, Xinyun Chen, Yuandong Tian, Jishen Zhao
Abstract
Task-specific fine-tuning on pre-trained transformers has achieved performance breakthroughs in multiple NLP tasks. Yet, as both computation and parameter size grows linearly with the number of sub-tasks, it is increasingly difficult to adopt such methods to the real world due to unrealistic memory and computation overhead on computing devices. Previous works on fine-tuning focus on reducing the growing parameter size to save storage cost by parameter sharing. However, compared to storage, the constraint of computation is a more critical issue with the fine-tuning models in modern computing environments. In this work, we propose LeTS, a framework that leverages both computation and parameter sharing across multiple tasks. Compared to traditional fine-tuning, LeTS proposes a novel neural architecture that contains a fixed pre-trained transformer model, plus learnable additive components for sub-tasks. The learnable components reuse the intermediate activations in the fixed pre-trained model, decoupling computation dependency. Differentiable neural architecture search is used to determine a task-specific computation sharing scheme, and a novel early stage pruning is applied to additive components for sparsity to achieve parameter sharing. Extensive experiments show that with 1.4% of extra parameters per task, LeTS reduces the computation by 49.5% on GLUE benchmarks with only 0.2% accuracy loss compared to full fine-tuning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 45c5d3b1-00a0-4a2a-9fc7-3c43f407a144Cited by top-tier papers12
- ReFT: Representation Finetuning for Language ModelsZhengxuan Wu, Aryaman Arora, Zheng Wang, Atticus Geiger et al.NeurIPS 2024 · 233 citations
- LLM-Adapters: An Adapter Family for Parameter-Efficient Fine-Tuning of Large Language ModelsZhiqiang Hu, Lei Wang, Yihuai Lan, Wanyu Xu et al.EMNLP 2023 · 200 citations
- Head2Toe: Utilizing Intermediate Representations for Better Transfer LearningUtku Evci, Vincent Dumoulin, Hugo Larochelle, Michael C. MozerICML 2022 · 103 citations
- PetS: A Unified Framework for Parameter-Efficient Transformers ServingZhe Zhou, Xuechao Wei, Jiejing Zhang, Guangyu SunUSENIX ATC 2022 · 41 citations
- Improved Representation Steering for Language ModelsZhengxuan Wu, Qinan Yu, Aryaman Arora, Christopher D. Manning et al.NeurIPS 2025 · 22 citations
Builds on5
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited DevicesZhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu et al.ACL 2020 · 660 citations
- UDapter: Language Adaptation for Truly Universal Dependency ParsingAhmet Üstün, Arianna Bisazza, Gosse Bouma, Gertjan van NoordEMNLP 2020 · 10 citations
- Parameter-Efficient Transfer Learning with Diff PruningDemi Guo, Alexander M. Rush, Yoon KimACL 2021
- FBNetV2: Differentiable Neural Architecture Search for Spatial and Channel DimensionsAlvin Wan, Xiaoliang Dai, Peizhao Zhang, Zijian He et al.CVPR 2020
Related papers
- Few-shot Task-agnostic Neural Architecture Search for Distilling Large Language ModelsDongkuan Xu, Subhabrata Mukherjee, Xiaodong Liu, Debadeepta Dey et al.NeurIPS 2022 · 21 citations
- Task Adaptive Parameter Sharing for Multi-Task LearningMatthew Wallingford, Hao Li, Alessandro Achille, Avinash Ravichandran et al.CVPR 2022 · 61 citations
- HyperPrompt: Prompt-based Task-Conditioning of TransformersYun He, Huaixiu Steven Zheng, Yi Tay, Jai Prakash Gupta et al.ICML 2022 · 110 citations
- Conditionally Adaptive Multi-Task Learning: Improving Transfer Learning in NLP Using Fewer Parameters & Less DataJonathan Pilault, Amine Elhattami, Christopher J. PalICLR 2021 · 105 citations
- TaskFusion: An Efficient Transfer Learning Architecture with Dual Delta Sparsity for Multi-Task Natural Language ProcessingZichen Fan, Qirui Zhang, Pierre Abillama, Sara Shoouri et al.ISCA 2023 · 15 citations
