Prompt Learning via Meta-Regularization
Jinyoung Park, Juyeon Ko, Hyunwoo J. Kim
Abstract
Pre-trained vision-language models have shown impres-sive success on various computer vision tasks with their zero-shot generalizability. Recently, prompt learning approaches have been explored to efficiently and effectively adapt the vision-language models to a variety of down-stream tasks. However, most existing prompt learning meth-ods suffer from task overfitting since the general knowl-edge of the pre-trained vision language models is forgot-ten while the prompts are finetuned on a small data set from a specific target task. To address this issue, we pro-pose a Prompt Meta-Regularization (ProMetaR) to im-prove the generalizability of prompt learning for vision-language models. Specifically, ProMetaR meta-learns both the regularizer and the soft prompts to harness the task-specific knowledge from the downstream tasks and task-agnostic general knowledge from the vision-language mod-els. Further, ProMetaR augments the task to gener-ate multiple virtual tasks to alleviate the meta-overfitting. In addition, we provide the analysis to comprehend how ProMetaR improves the generalizability of prompt tuning in the perspective of the gradient alignment. Our exten-sive experiments demonstrate that our ProMetaR improves the generalizability of conventional prompt learning meth-ods under base-to-base/base-to-new and domain general-ization settings. The code of ProMetaR is available at https://github.com/mlvlab/ProMetaR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 09dfa38f-f932-4edd-ae96-6f626bb25e4aCited by top-tier papers20
- X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action ModelJinliang Zheng, Jianxiong Li, Zhihao Wang, Dongxiu Liu et al.ICLR 2026 · 335 citations
- VaMP: Variational Multi-Modal Prompt Learning for Vision-Language ModelsSilin Cheng, Kai HanNeurIPS 2025 · 7 citations
- Multimodal Prompt Alignment for Facial Expression RecognitionFuyan Ma, Yiran He, Bin Sun, Shutao LiICCV 2025 · 5 citations
- Visual Diversity and Region-aware Prompt Learning for Zero-shot HOI DetectionChanhyeong Yang, Taehoon Song, Jihwan Park, Hyunwoo J. KimNeurIPS 2025 · 5 citations
- Modality Alignment across Trees on Heterogeneous Hyperbolic ManifoldsWei Wu, Xiaomeng Fan, Yuwei Wu, Zhi Gao et al.ICLR 2026 · 3 citations
Builds on31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
Related papers
- Gradient-Regulated Meta-Prompt Learning for Generalizable Vision-Language ModelsJuncheng Li, Minghe Gao, Longhui Wei, Siliang Tang et al.ICCV 2023 · 34 citations
- Debiased Fine-Tuning for Vision-Language Models by Prompt RegularizationBeier Zhu, Yulei Niu, Saeil Lee, Minhoe Hur et al.AAAI 2023 · 34 citations
- A Similarity Paradigm Through Textual Regularization Without ForgettingFangming Cui, Jan Fong, Rongfei Zeng, Xinmei Tian et al.AAAI 2025 · 4 citations
- Consistency-guided Prompt Learning for Vision-Language ModelsShuvendu Roy, Ali EtemadICLR 2024 · 102 citations
- Distribution-Aware Prompt Tuning for Vision-Language ModelsEulrang Cho, Jooyeon Kim, Hyunwoo J. KimICCV 2023 · 54 citations
