Amend to Alignment: Decoupled Prompt Tuning for Mitigating Spurious Correlation in Vision-Language Models
Jie Zhang, Xiaosong Ma, Song Guo, Peng Li, Wenchao Xu, Xueyang Tang, Zicong Hong
Abstract
Fine-tuning the learnable prompt for a pre-trained vision-language model (VLM), such as CLIP, has demonstrated exceptional efficiency in adapting to a broad range of downstream tasks. Existing prompt tuning methods for VLMs do not distinguish spurious features introduced by biased training data from invariant features, and employ a uniform alignment process when adapting to unseen target domains. This can impair the cross-modal feature alignment when the test data significantly deviate from the distribution of the training data, resulting in a poor out-ofdistribution (OOD) generalization performance. In this paper, we reveal that the prompt tuning failure in such OOD scenarios can be attribute to the undesired alignment between the textual and the spurious feature. As a solution, we propose CoOPood, a fine-grained prompt tuning method that can discern the causal features and deliberately align the text modality with the invariant feature. Specifically, we design two independent contrastive phases using two lightweight projection layers during the alignment, each with different objectives: 1) pulling the text embedding closer to the invariant image embedding and 2) pushing the text embedding away from the spurious image embedding. We have illustrated that CoOPood can serve as a general framework for VLMs and can be seamlessly integrated with existing prompt tuning methods. Extensive experiments on various OOD datasets demonstrate the performance superiority over state-of-the-art methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- DiMPLE - Disentangled Multi-Modal Prompt Learning: Enhancing Out-of-Distribution Alignment with Invariant and Spurious Feature SeparationUmaima Rahman, Mohammad Yaqub, Dwarikanath MahapatraICCV 2025 · 1 citation
- SAGE: Spuriousness-Aware Guided Prompt Exploration for Mitigating Multimodal BiasWenqian Ye, Di Wang, Guangtao Zheng, Bohan Liu et al.AAAI 2026
- Supporting Multimodal Intermediate Fusion with Informatic Constraint and Distribution CoherenceYi Li, Fei Song, Changwen Zheng, Jiangmeng LiICLR 2026
- CoCoA-Mix: Confusion-and-Confidence-Aware Mixture Model for Context OptimizationDasol Hong, Wooju Lee, Hyun MyungICML 2025
- Density-Aware Translation of Spurious Correlations in Zero-Shot VLMsAfsaneh Hasanebrahimi, Hanxun Huang, Christopher Leckie, Sarah ErfaniICML 2026
Builds on40
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
Related papers
- MaPLe: Multi-modal Prompt LearningMuhammad Uzair Khattak, Hanoona Abdul Rasheed, Muhammad Maaz, Salman H. Khan et al.CVPR 2023
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 1,438 citations
- Concept-Guided Prompt Learning for Generalization in Vision-Language ModelsYi Zhang, Ce Zhang, Ke Yu, Yushun Tang et al.AAAI 2024 · 37 citations
- FA: Forced Prompt Learning of Vision-Language Models for Out-of-Distribution DetectionXinhua Lu, Runhe Lai, Yanqi Wu, Kanghao Chen et al.ICCV 2025 · 1 citation
- COMMA: Co-articulated Multi-Modal LearningLianyu Hu, Liqing Gao, Zekang Liu, Chi-Man Pun et al.AAAI 2024 · 7 citations
