Prompt-Robust Vision-Language Models via Meta-Finetuning
Haohui Liang, Runlin Huang, Yingjun Du, Yujia Hu, Weifeng Su, Cees G. M. Snoek
Abstract
Vision-language models (VLMs) have demonstrated remarkable generalization across diverse tasks by leveraging large-scale image-text pretraining. However, their performance is notoriously unstable under variations in natural language prompts, posing a considerable challenge for reliable real-world deployment. To address this prompt sensitivity, we propose Promise, a meta-learning framework for prompt-Robust vision-language models via meta-finetuning, which explicitly learns to generalize across diverse prompt formulations. Our method operates in a dual-loop meta-finetuning setting: the inner loop adapts token embeddings based on a set of varied prompts, while the outer loop optimizes for generalization on unseen prompt variants. To further improve robustness, we introduce an adaptive prompt weighting mechanism that dynamically emphasizes more generalizable prompts and a token-specific learning rate module that fine-tunes individual prompt tokens based on contextual importance. We further establish that Promise’s weighted and preconditioned inner update provably (i) yields a one-step decrease of the outer empirical risk together with a contraction of across-prompt sensitivity, and (ii) tightens a data-dependent generalization bound evaluated at the post-inner initialization. Across 15 benchmarks spanning base-to-novel generalization, cross-dataset transfer, and domain shift, our approach consistently reduces prompt sensitivity and improves performance stability over existing prompt learning methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b5388c87-7d67-4108-b740-2bcfed51879fCited by top-tier papers1
Ask how each one uses itBuilds on19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath et al.ICCV 2021 · 2,294 citations
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 1,438 citations
- Self-regulating Prompts: Foundational Model Adaptation without ForgettingMuhammad Uzair Khattak, Syed Talal Wasim, Muzammal Naseer, Salman Khan et al.ICCV 2023 · 365 citations
Related papers
- Gradient-Regulated Meta-Prompt Learning for Generalizable Vision-Language ModelsJuncheng Li, Minghe Gao, Longhui Wei, Siliang Tang et al.ICCV 2023 · 34 citations
- One Prompt Word is Enough to Boost Adversarial Robustness for Pre-Trained Vision-Language ModelsLin Li, Haoyan Guan, Jianing Qiu, Michael W. SpratlingCVPR 2024
- Learning Robust Vision-Language Models from Natural Latent SpacesZhangyun Wang, Ni Ding, Aniket MahantiNeurIPS 2025 · 3 citations
- A Retrospect to Multi-prompt Learning across Vision and LanguageZiliang Chen, Xin Huang, Quanlong Guan, Liang Lin et al.ICCV 2023 · 12 citations
- Prompt Learning via Meta-RegularizationJinyoung Park, Juyeon Ko, Hyunwoo J. KimCVPR 2024 · 17 citations
