A Retrospect to Multi-prompt Learning across Vision and Language
Ziliang Chen, Xin Huang, Quanlong Guan, Liang Lin, Weiqi Luo
摘要
The vision community is undergoing the unprecedented progress with the emergence of Vision-Language Pretraining Models (VLMs). Prompt learning plays as the holy grail of accessing VLMs since it enables their fast adaptation to downstream tasks with limited resources. Whereas existing researches milling around single-prompt paradigms, rarely investigate the technical potential behind their multi-prompt learning counterparts. This paper aims to provide a principled retrospect for vision-language multi-prompt learning. We extend the recent constant modality gap phenomenon to learnable prompts and then, justify the superiority of vision-language transfer with multi-prompt augmentation, empirically and theoretically. In terms of this observation, we propose an Energy-based Multi-prompt Learning (EMPL) to generate multiple prompt embeddings by drawing instances from an energy-based distribution, which is implicitly defined by VLMs. So our EMPL is not only parameter-efficient but also rigorously lead to the balance between in-domain and out-of-domain open-vocabulary generalization. Comprehensive experiments have been conducted to justify our claims and the excellence of EMPL.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Federated Text-driven Prompt Generation for Vision-Language ModelsChen Qiu, Xingyu Li, Chaithanya Kumar Mummadi, Madan Ravi Ganesh 等ICLR 2024 · 被引用 33 次
- Causality-Guided Prompt Learning for Vision-Language Models via Visual GranulationMengyu Gao, Qiulei DongICCV 2025 · 被引用 2 次
- Prompt Tuning In a Compact Attribute SpaceShiyu Hou, Tianfei Zhou, Shuai Zhang, Ye Yuan 等AAAI 2025 · 被引用 1 次
- A Causal Marriage between VLM and IRM from Understanding to ReasoningZiliang Chen, Tianang Xiao, jusheng zhang, Yongsen Zheng 等CVPR 2026
- Vocabulary Scaling Law: Tuning Open-vocabulary Predictors for Their OpennessZiliang Chen, Yulu Li, Liangda Fang, jusheng zhang 等CVPR 2026
它引用的顶会 Paper24
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsAlexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam 等ICML 2022 · 被引用 4,691 次
相关 Paper
- Advancing Prompt Learning through an External LayerFangming Cui, Xun Yang, Chao Wu, Liang Xiao 等ACM MM 2024 · 被引用 3 次
- Prompt-Robust Vision-Language Models via Meta-FinetuningHaohui Liang, Runlin Huang, Yingjun Du, Yujia Hu 等ICLR 2026
- Distribution-Aware Prompt Tuning for Vision-Language ModelsEulrang Cho, Jooyeon Kim, Hyunwoo J. KimICCV 2023 · 被引用 54 次
- MaPLe: Multi-modal Prompt LearningMuhammad Uzair Khattak, Hanoona Abdul Rasheed, Muhammad Maaz, Salman H. Khan 等CVPR 2023
- Aggregate-and-Adapt Natural Language Prompts for Downstream Generalization of CLIPChen Huang, Skyler Seto, Samira Abnar, David Grangier 等NeurIPS 2024 · 被引用 8 次
