Prototype-based HyperAdapter for Sample-Efficient Multi-task Tuning
Hao Zhao, Jie Fu, Zhaofeng He
Abstract
Parameter-efficient fine-tuning has shown its effectiveness in adapting the pre-trained language models to downstream tasks while only updating a small number of parameters. Despite the success, most existing methods independently adapt to each task without considering knowledge transfer between tasks and are limited to low-data regimes. To overcome this issue, we propose Prototype-based Hyper-Adapter (PHA), a novel framework built on the adapter-tuning and hypernetwork. It introduces an instance-dense retriever and a prototypical hypernetwork to generate the conditional modules in a sample-efficient manner. This leads to comparable performance improvements against existing Parameter-efficient fine-tuning methods on multi-task learning and few-shot transfer learning. More importantly, when the available data size gets smaller, our method outperforms other strong baselines by a large margin. Based on our extensive empirical experiments across various datasets, we demonstrate that PHA strikes a better trade-off between trainable parameters, accuracy on stream tasks, and sample efficiency. Our code is publicly available at https://github.com/Bumble666/PHA
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 51473fb4-c69d-42d2-b594-c93044cdc6faCited by top-tier papers2
- Beyond Adapter Retrieval: Latent Geometry-Preserving Composition via Sparse Task ProjectionPengfei Jin, Peng Shu, Sifan Song, Sekeun Kim et al.AAAI 2026 · 2 citations
- HyperMoE: Towards Better Mixture of Experts via Transferring Among ExpertsHao Zhao, Zihan Qiu, Huijia Wu, Zili Wang et al.ACL 2024
Builds on21
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach et al.ICLR 2022 · 1,976 citations
- Compacter: Efficient Low-Rank Hypercomplex Adapter LayersRabeeh Karimi Mahabadi, James Henderson, Sebastian RuderNeurIPS 2021 · 700 citations
Related papers
- Parameter-efficient Multi-task Fine-tuning for Transformers via Shared HypernetworksRabeeh Karimi Mahabadi, Sebastian Ruder, Mostafa Dehghani, James HendersonACL 2021
- MTLoRA: A Low-Rank Adaptation Approach for Efficient Multi-Task LearningAhmed Agiza, Marina Neseem, Sherief RedaCVPR 2024
- VMT-Adapter: Parameter-Efficient Transfer Learning for Multi-Task Dense Scene UnderstandingYi Xin, Junlong Du, Qiang Wang, Zhiwen Lin et al.AAAI 2024 · 94 citations
- HyperPrompt: Prompt-based Task-Conditioning of TransformersYun He, Huaixiu Steven Zheng, Yi Tay, Jai Prakash Gupta et al.ICML 2022 · 110 citations
- ATTEMPT: Parameter-Efficient Multi-task Tuning via Attentional Mixtures of Soft PromptsAkari Asai, Mohammadreza Salehi, Matthew E. Peters, Hannaneh HajishirziEMNLP 2022 · 55 citations
