Gradient-Regulated Meta-Prompt Learning for Generalizable Vision-Language Models
Juncheng Li, Minghe Gao, Longhui Wei, Siliang Tang, Wenqiao Zhang, Mengze Li, Wei Ji, Qi Tian, Tat-Seng Chua, Yueting Zhuang
Abstract
Prompt tuning, a recently emerging paradigm, enables the powerful vision-language pre-training models to adapt to downstream tasks in a parameter- and data- efficient way, by learning the "soft prompts" to condition frozen pretraining models. Though effective, it is particularly problematic in the few-shot scenario, where prompt tuning performance is sensitive to the initialization and requires a time-consuming process to find a good initialization, thus restricting the fast adaptation ability of the pre-training models. In addition, prompt tuning could undermine the generalizability of the pre-training models, because the learnable prompt tokens are easy to overfit to the limited training samples. To address these issues, we introduce a novel Gradient-RegulAted Meta-prompt learning (GRAM) framework that jointly meta-learns an efficient soft prompt initialization for better adaptation and a lightweight gradient regulating function for strong cross-domain generalizability in a meta-learning paradigm using only the unlabeled image-text pre-training data. Rather than designing a specific prompt tuning method, our GRAM can be easily incorporated into various prompt tuning methods in a model-agnostic way, and comprehensive experiments show that GRAM brings about consistent improvement for them in several settings (i.e., few-shot learning, cross-domain generalization, cross-dataset generalization, etc.) over 11 datasets. Further, experiments show that GRAM enables the orthogonal methods of textual and visual prompt tuning to work in a mutually-enhanced way, offering better generalizability beyond the uni-modal prompt tuning methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 64dc8654-da2d-4006-98ee-ad02c6f4d0f3Cited by top-tier papers10
- TCP: Textual-Based Class-Aware Prompt Tuning for Visual-Language ModelHantao Yao, Rui Zhang, Changsheng XuCVPR 2024 · 46 citations
- AWT: Transferring Vision-Language Models via Augmentation, Weighting, and TransportationYuhan Zhu, Yuyang Ji, Zhiyu Zhao, Gangshan Wu et al.NeurIPS 2024 · 45 citations
- Data Shunt: Collaboration of Small and Large Models for Lower Costs and Better PerformanceDong Chen, Yueting Zhuang, Shuo Zhang, Jinfeng Liu et al.AAAI 2024 · 32 citations
- Unified Generative and Discriminative Training for Multi-modal Large Language ModelsWei Chow, Juncheng Li, Qifan Yu, Kaihang Pan et al.NeurIPS 2024 · 19 citations
- Few-Shot Image Quality Assessment via Adaptation of Vision-Language ModelsXudong Li, Zihao Huang, Yan Zhang, Yunhang Shen et al.ICCV 2025 · 2 citations
Builds on19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath et al.ICCV 2021 · 2,294 citations
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 1,438 citations
- Prompt-aligned Gradient for Prompt TuningBeier Zhu, Yulei Niu, Yucheng Han, Yue Wu et al.ICCV 2023 · 475 citations
Related papers
- Prompt Learning via Meta-RegularizationJinyoung Park, Juyeon Ko, Hyunwoo J. KimCVPR 2024 · 17 citations
- PPT: Pre-trained Prompt Tuning for Few-shot LearningYuxian Gu, Xu Han, Zhiyuan Liu, Minlie HuangACL 2022
- Learning to Initialize: Can Meta Learning Improve Cross-task Generalization in Prompt Tuning?Chengwei Qin, Shafiq R. Joty, Qian Li, Ruochen ZhaoACL 2023 · 8 citations
- Prompt-Robust Vision-Language Models via Meta-FinetuningHaohui Liang, Runlin Huang, Yingjun Du, Yujia Hu et al.ICLR 2026
- Generalizing Vision-Language Models with Dedicated Prompt GuidanceXinyao Li, Yinjie Min, Hongbo Chen, Zhekai Du et al.AAAI 2026
