Learning Invariant Causal Mechanism from Vision-Language Models
Zeen Song, Siyu Zhao, Xingyu Zhang, Jiangmeng Li, Changwen Zheng, Wenwen Qiang
摘要
Contrastive Language-Image Pretraining (CLIP) has achieved remarkable success, but its performance can degrade when fine-tuned in out-ofdistribution (OOD) scenarios. We model the prediction process using a Structural Causal Model (SCM) and show that the causal mechanism involving both invariant and variant factors in training environments differs from that in test environments. In contrast, the causal mechanism with solely invariant factors remains consistent across environments. We theoretically prove the existence of a linear mapping from CLIP embeddings to invariant factors, which can be estimated using interventional data. Additionally, we provide a condition to guarantee low OOD risk of the invariant predictor. Based on these insights, we propose the Invariant Causal Mechanism of CLIP (CLIP-ICM) framework. CLIP-ICM involves collecting interventional data, estimating a linear projection matrix, and making predictions within the invariant subspace. Experiments on several OOD datasets show that CLIP-ICM significantly improves the performance of CLIP. Our method offers a simple but powerful enhancement, boosting the reliability of CLIP in real-world applications. The source code is available at https://github.com/ZeenSong/CLIP-ICM .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Enhancing Reward Models for High-Quality Image Generation: Beyond Text-Image AlignmentYing Ba, Tianyu Zhang, Yalong Bai, Wenyi Mo 等ICCV 2025 · 被引用 13 次
- Progressive Cross-Modal Causal Intervention for Long-Term Action RecognitionShaowu Xu, Xibin Jia, Chao Fan, Junyu Gao 等CVPR 2026
- TMAE: Learning Targeted Multi-Agent Exploration via Causal InferenceChuxiong Sun, Dunqi Yao, Rui Wang, Wenwen Qiang 等AAAI 2026
- Learning Invariant Modality Representation for Robust Multimodal Learning from a Causal Inference PerspectiveSijie Mai, Shiqin HanACL 2026
- Cross-Modal Dual-Causal Learning for Long-Term Action RecognitionShaowu Xu, Xibin Jia, Junyu Gao, Qianmei Sun 等ACM MM 2025
它引用的顶会 Paper28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath 等ICCV 2021 · 被引用 2,294 次
相关 Paper
- A Causal Marriage between VLM and IRM from Understanding to ReasoningZiliang Chen, Tianang Xiao, jusheng zhang, Yongsen Zheng 等CVPR 2026
- Amend to Alignment: Decoupled Prompt Tuning for Mitigating Spurious Correlation in Vision-Language ModelsJie Zhang, Xiaosong Ma, Song Guo, Peng Li 等ICML 2024 · 被引用 10 次
- Whitened CLIP as a Likelihood Surrogate of Images and CaptionsRoy Betser, Meir Yossef Levi, Guy GilboaICML 2025
- CLIP Behaves like a Bag-of-Words Model Cross-modally but not Uni-modallyDarina Koishigarina, Arnas Uselis, Seong Joon OhICLR 2026 · 被引用 33 次
- A Closer Look at the Robustness of Contrastive Language-Image Pre-Training (CLIP)Weijie Tu, Weijian Deng, Tom GedeonNeurIPS 2023 · 被引用 74 次
