Task Adaptive Parameter Sharing for Multi-Task Learning
Matthew Wallingford, Hao Li, Alessandro Achille, Avinash Ravichandran, Charless C. Fowlkes, Rahul Bhotika, Stefano Soatto
摘要
Adapting pre-trained models with broad capabilities has become standard practice for learning a wide range of downstream tasks. The typical approach of fine-tuning different models for each task is performant, but incurs a substantial memory cost. To efficiently learn multiple down-stream tasks we introduce Task Adaptive Parameter Sharing (TAPS), a simple method for tuning a base model to a new task by adaptively modifying a small, task-specific subset of layers. This enables multi-task learning while minimizing the resources used and avoids catastrophic forgetting and competition between tasks. TAPS solves a joint optimization problem which determines both the layers that are shared with the base model and the value of the task-specific weights. Further, a sparsity penalty on the number of active layers promotes weight sharing with the base model. Compared to other methods, TAPS retains a high accuracy on the target tasks while still introducing only a small number of task-specific parameters. Moreover, TAPS is agnostic to the particular architecture used and requires only minor changes to the training scheme. We evaluate our method on a suite of fine-tuning tasks and architectures (ResNet, DenseNet, ViT) and show that it achieves state-of-the-art performance while being simple to implement.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Matryoshka Representation LearningAditya Kusupati, Gantavya Bhatt, Aniket Rege, Matthew Wallingford 等NeurIPS 2022 · 被引用 364 次
- Towards Modular LLMs by Building and Reusing a Library of LoRAsOleksiy Ostapenko, Zhan Su, Edoardo M. Ponti, Laurent Charlin 等ICML 2024 · 被引用 70 次
- Neural Network Architecture Beyond Width and DepthShijun Zhang, Zuowei Shen, Haizhao YangNeurIPS 2022 · 被引用 25 次
- M3Net: Multimodal Multi-task Learning for 3D Detection, Segmentation, and Occupancy Prediction in Autonomous DrivingXuesong Chen, Shaoshuai Shi, Tao Ma, Jingqiu Zhou 等AAAI 2025 · 被引用 14 次
- Exploring Training on Heterogeneous Data with Mixture of Low-rank AdaptersYuhang Zhou, Zihua Zhao, Siyuan Du, Haolin Li 等ICML 2024 · 被引用 11 次
它引用的顶会 Paper4
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Moment Matching for Multi-Source Domain AdaptationXingchao Peng, Qinxun Bai, Xide Xia, Zijun Huang 等ICCV 2019 · 被引用 2,239 次
- AdaShare: Learning What To Share For Efficient Deep Multi-Task LearningXimeng Sun, Rameswar Panda, Rogério Feris, Kate SaenkoNeurIPS 2020 · 被引用 337 次
- Budget-Aware Adapters for Multi-Domain LearningRodrigo Ferreira Berriel, Stéphane Lathuilière, Moin Nabi, Tassilo Klein 等ICCV 2019 · 被引用 45 次
相关 Paper
- Your representations are in the network: composable and parallel adaptation for large scale modelsYonatan Dukler, Alessandro Achille, Hao Yang, Varsha Vivek 等NeurIPS 2023 · 被引用 4 次
- Fine-tuning Image Transformers using Learnable MemoryMark Sandler, Andrey Zhmoginov, Max Vladymyrov, Andrew JacksonCVPR 2022 · 被引用 51 次
- Polyhistor: Parameter-Efficient Multi-Task Adaptation for Dense Vision TasksYen-Cheng Liu, Chih-Yao Ma, Junjiao Tian, Zijian He 等NeurIPS 2022 · 被引用 79 次
- Learn-to-Share: A Hardware-friendly Transfer Learning Framework Exploiting Computation and Parameter SharingCheng Fu, Hanxian Huang, Xinyun Chen, Yuandong Tian 等ICML 2021 · 被引用 28 次
- Memory Efficient Continual Learning with TransformersBeyza Ermis, Giovanni Zappella, Martin Wistuba, Aditya Rawal 等NeurIPS 2022 · 被引用 75 次
