Multitask Pre-training of Modular Prompt for Chinese Few-Shot Learning
Tianxiang Sun, Zhengfu He, Qin Zhu, Xipeng Qiu, Xuanjing Huang
Abstract
Prompt tuning is a parameter-efficient approach to adapting pre-trained language models to downstream tasks. Although prompt tuning has been shown to match the performance of full model tuning when training data is sufficient, it tends to struggle in few-shot learning settings. In this paper, we present Multi-task Pre-trained Modular Prompt (MP 2 ) to boost prompt tuning for few-shot learning. MP 2 is a set of combinable prompts pre-trained on 38 Chinese tasks. On downstream tasks, the pre-trained prompts are selectively activated and combined, leading to strong compositional generalization to unseen tasks. To bridge the gap between pre-training and fine-tuning, we formulate upstream and downstream tasks into a unified machine reading comprehension task. Extensive experiments under two learning paradigms, i.e., gradient descent and black-box tuning, show that MP 2 significantly outperforms prompt tuning, full model tuning, and prior prompt pretraining methods in few-shot settings. In addition, we demonstrate that MP 2 can achieve surprisingly fast and strong adaptation to downstream tasks by merely learning 8 parameters to combine the pre-trained modular prompts.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ee58e9cd-8f4f-47c0-9f53-5bcc3b8cef51Cited by top-tier papers5
- When MOE Meets LLMs: Parameter Efficient Fine-tuning for Multi-task Medical ApplicationsQidong Liu, Xian Wu, Xiangyu Zhao, Yuanshao Zhu et al.SIGIR 2024 · 89 citations
- Decomposed Prompt Decision Transformer for Efficient Unseen Task GeneralizationHongling Zheng, Li Shen, Yong Luo, Tongliang Liu et al.NeurIPS 2024 · 13 citations
- Adaptive Stochastic Gradient Algorithm for Black-box Multi-Objective LearningFeiyang Ye, Yueming Lyu, Xuehao Wang, Yu Zhang et al.ICLR 2024 · 5 citations
- Sharpness-Aware Black-Box OptimizationFeiyang Ye, Yueming Lyu, Xuehao Wang, Masashi Sugiyama et al.ICLR 2025
- MeteoRA: Multiple-tasks Embedded LoRA for Large Language ModelsJingwei Xu, Junyu Lai, Yunpeng HuangICLR 2025
Builds on12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Towards a Unified View of Parameter-Efficient Transfer LearningJunxian He, Chunting Zhou, Xuezhe Ma, Taylor Berg-Kirkpatrick et al.ICLR 2022 · 1,182 citations
Related papers
- PPT: Pre-trained Prompt Tuning for Few-shot LearningYuxian Gu, Xu Han, Zhiyuan Liu, Minlie HuangACL 2022
- BBTv2: Towards a Gradient-Free Future with Large Language ModelsTianxiang Sun, Zhengfu He, Hong Qian, Yunhua Zhou et al.EMNLP 2022 · 37 citations
- Gradient-Regulated Meta-Prompt Learning for Generalizable Vision-Language ModelsJuncheng Li, Minghe Gao, Longhui Wei, Siliang Tang et al.ICCV 2023 · 34 citations
- ATTEMPT: Parameter-Efficient Multi-task Tuning via Attentional Mixtures of Soft PromptsAkari Asai, Mohammadreza Salehi, Matthew E. Peters, Hannaneh HajishirziEMNLP 2022 · 55 citations
- Multitask Prompt Tuning Enables Parameter-Efficient Transfer LearningZhen Wang, Rameswar Panda, Leonid Karlinsky, Rogério Feris et al.ICLR 2023 · 30 citations
