Tuning Multi-mode Token-level Prompt Alignment across Modalities
Dongsheng Wang, Miaoge Li, Xinyang Liu, Mingsheng Xu, Bo Chen, Hanwang Zhang
摘要
Advancements in prompt tuning of vision-language models have underscored their potential in enhancing open-world visual concept comprehension. However, prior works only primarily focus on single-mode (only one prompt for each modality) and holistic level (image or sentence) semantic alignment, which fails to capture the sample diversity, leading to sub-optimal prompt discovery. To address the limitation, we propose a multi-mode token-level tuning framework that leverages the optimal transportation to learn and align a set of prompt tokens across modalities. Specifically, we rely on two essential factors: 1) multi-mode prompts discovery, which guarantees diverse semantic representations, and 2) token-level alignment, which helps explore fine-grained similarity. Consequently, the similarity can be calculated as a hierarchical transportation problem between the modality-specific sets. Extensive experiments on popular image recognition benchmarks show the superior generalization and few-shot abilities of our approach. The qualitative analysis demonstrates that the learned prompt tokens have the ability to capture diverse visual concepts. The code is available at https://github.com/wds2014/ALIGN .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- Enhancing Minority Classes by Mixing: An Adaptative Optimal Transport Approach for Long-tailed ClassificationJintong Gao, He Zhao, Zhuo Li, Dandan GuoNeurIPS 2023 · 被引用 64 次
- AWT: Transferring Vision-Language Models via Augmentation, Weighting, and TransportationYuhan Zhu, Yuyang Ji, Zhiyu Zhao, Gangshan Wu 等NeurIPS 2024 · 被引用 45 次
- ArGue: Attribute-Guided Prompt Tuning for Vision-Language ModelsXinyu Tian, Shu Zou, Zhaoyuan Yang, Jing ZhangCVPR 2024 · 被引用 29 次
- Enhancing CLIP Robustness via Cross-Modality AlignmentXingyu Zhu, Beier Zhu, Shuo Wang, Kesen Zhao 等NeurIPS 2025 · 被引用 17 次
- Look Carefully: Adaptive Visual Reinforcements in Multimodal Large Language Models for Hallucination MitigationXingyu Zhu, Kesen Zhao, Liang Yi, Shuo Wang 等ICLR 2026 · 被引用 9 次
它引用的顶会 Paper23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath 等ICCV 2021 · 被引用 2,294 次
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 被引用 1,438 次
相关 Paper
- PLOT: Prompt Learning with Optimal Transport for Vision-Language ModelsGuangyi Chen, Weiran Yao, Xiangchen Song, Xinyue Li 等ICLR 2023 · 被引用 22 次
- Distribution-Aware Prompt Tuning for Vision-Language ModelsEulrang Cho, Jooyeon Kim, Hyunwoo J. KimICCV 2023 · 被引用 54 次
- Prompting Multi-Modal Image Segmentation with Semantic GroupingQibin HeAAAI 2024 · 被引用 21 次
- Noise-Aware Few-Shot Learning through Bi-directional Multi-View Prompt AlignmentLu Niu, Cheng XueCVPR 2026
- CoPL: Contextual Prompt Learning for Vision-Language UnderstandingKoustava Goswami, Srikrishna Karanam, Prateksha Udhayanan, K. J. Joseph 等AAAI 2024 · 被引用 20 次
