Understanding the Robustness of Multi-modal Contrastive Learning to Distribution Shift
Yihao Xue, Siddharth Joshi, Dang Nguyen, Baharan Mirzasoleiman
Abstract
Recently, multimodal contrastive learning (MMCL) approaches, such as CLIP (Radford et al., 2021) , have achieved a remarkable success in learning representations that are robust against distribution shift and generalize to new domains. Despite the empirical success, the mechanism behind learning such generalizable representations is not understood. In this work, we rigorously analyze this problem and uncover two mechanisms behind MMCL's robustness: intra-class contrasting, which allows the model to learn features with a high variance, and inter-class feature sharing, where annotated details in one class help learning other classes better. Both mechanisms prevent spurious features that are over-represented in the training data to overshadow the generalizable core features. This yields superior zeroshot classification accuracy under distribution shift. Furthermore, we theoretically demonstrate the benefits of using rich captions on robustness and explore the effect of annotating different types of details in the captions. We validate our theoretical findings through experiments, including a well-designed synthetic experiment and an experiment involving training CLIP models on MSCOCO (Lin et al., 2014)/Conceptual Captions (Sharma et al., 2018) and evaluating them on shifted ImageNets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 951ff24c-5891-424e-89d7-a602b2677f2bCited by top-tier papers12
- A Sober Look at the Robustness of CLIPs to Spurious FeaturesQizhou Wang, Yong Lin, Yongqiang Chen, Ludwig Schmidt et al.NeurIPS 2024 · 46 citations
- Temporal-Frequency Masked Autoencoders for Time Series Anomaly DetectionYuchen Fang, Jiandong Xie, Yan Zhao, Lu Chen et al.ICDE 2024 · 45 citations
- On the Comparison between Multi-modal and Single-modal Contrastive LearningWei Huang, Andi Han, Yongqiang Chen, Yuan Cao et al.NeurIPS 2024 · 26 citations
- Few-shot Adaptation to Distribution Shifts By Mixing Source and Target EmbeddingsYihao Xue, Ali Payani, Yu Yang, Baharan MirzasoleimanICML 2024 · 4 citations
- Toward Enhancing Representation Learning in Federated Multi-Task SettingsMehdi Setayesh, Mahdi Beitollahi, Yasser H. Khalil, Hongliang LiICLR 2026 · 2 citations
Builds on30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong et al.NeurIPS 2020 · 3,935 citations
Related papers
- Understanding Transferable Representation Learning and Zero-shot Transfer in CLIPZixiang Chen, Yihe Deng, Yuanzhi Li, Quanquan GuICLR 2024 · 21 citations
- Interpreting and Analysing CLIP's Zero-Shot Image Classification via Mutual KnowledgeFawaz Sammani, Nikos DeligiannisNeurIPS 2024 · 15 citations
- Data Determines Distributional Robustness in Contrastive Language Image Pre-training (CLIP)Alex Fang, Gabriel Ilharco, Mitchell Wortsman, Yuhao Wan et al.ICML 2022 · 183 citations
- When and How Does CLIP Enable Domain and Compositional Generalization?Elias Kempf, Simon Schrodi, Max Argus, Thomas BroxICML 2025
- Beyond DAGs: A Latent Partial Causal Model for Multimodal LearningYuhang Liu, Zhen Zhang, Dong Gong, Erdun Gao et al.ICLR 2026 · 9 citations
