A Sober Look at the Robustness of CLIPs to Spurious Features
Qizhou Wang, Yong Lin, Yongqiang Chen, Ludwig Schmidt, Bo Han, Tong Zhang
Abstract
Large vision language models, such as CLIP, demonstrate impressive robustness to spurious features than single-modal models trained on ImageNet. However, existing test datasets are typically curated based on ImageNet-trained models, which aim to capture the spurious features inherited in ImageNet. Benchmarking CLIP models based on the ImageNet-oriented spurious features may not be sufficient to reflect the extent to which CLIP models are robust to spurious correlations within CLIP training data, e.g., LAION. To this end, we craft a new challenging dataset named CounterAnimal designed to reveal the reliance of CLIP models on realistic spurious features. Specifically, we split animal photos into groups according to the backgrounds, and then identify a pair of groups for each class where a CLIP model shows high-performance drops across the two groups. Our evaluations show that the spurious features captured by CounterAnimal are generically learned by CLIP models with different backbones and pre-train data, yet have limited influence for ImageNet models. We provide theoretical insights that the CLIP objective cannot offer additional robustness. Furthermore, we also re-evaluate strategies such as scaling up parameters and high-quality pre-trained data. We find that they still help mitigate the spurious features, providing a promising path for future developments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eb44888a-1cd2-485b-8696-d5c54166afc7Cited by top-tier papers15
- Web Artifact Attacks Disrupt Vision Language ModelsMaan Qraitem, Piotr Teterwak, Kate Saenko, Bryan A. PlummerICCV 2025 · 4 citations
- H-SPLID: HSIC-based Saliency Preserving Latent Information DecompositionLukas Miklautz, Chengzhi Shi, Andrii Shkabrii, Theodoros-Thirimachos Davarakis et al.NeurIPS 2025 · 2 citations
- Less Precise Can Be More Reliable: A Systematic Evaluation of Quantization’s Impact on VLMs Beyond AccuracyAymen Bouguerra, Daniel Vasquez, Alexandra Gomez-Villa, Chokri Mraidha et al.ICML 2026 · 2 citations
- Concept Regions Matter: Benchmarking CLIP with a New Cluster-Importance ApproachAishwarya Agarwal, Srikrishna Karanam, Vineet GandhiCVPR 2026 · 2 citations
- Is Less More? Exploring Token Condensation as Training-Free Test-Time AdaptationZixin Wang, Dong Gong, Sen Wang, Zi Huang et al.ICCV 2025 · 1 citation
Builds on37
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
- Distributionally Robust Neural NetworksShiori Sagawa, Pang Wei Koh, Tatsunori B. Hashimoto, Percy LiangICLR 2020 · 1,578 citations
Related papers
- Does CLIP's generalization performance mainly stem from high train-test similarity?Prasanna Mayilvahanan, Thaddäus Wiedemer, Evgenia Rusak, Matthias Bethge et al.ICLR 2024 · 43 citations
- African or European Swallow? Benchmarking Large Vision-Language Models for Fine-Grained Object ClassificationGregor Geigle, Radu Timofte, Goran GlavasEMNLP 2024 · 2 citations
- Think Twice: Test-Time Reasoning for Robust CLIP Zero-Shot ClassificationShenyu Lu, Zhaoying Pan, Xiaoqian WangICCV 2025 · 1 citation
- SAGE: Spuriousness-Aware Guided Prompt Exploration for Mitigating Multimodal BiasWenqian Ye, Di Wang, Guangtao Zheng, Bohan Liu et al.AAAI 2026
- Quality Not Quantity: On the Interaction between Dataset Design and Robustness of CLIPThao Nguyen, Gabriel Ilharco, Mitchell Wortsman, Sewoong Oh et al.NeurIPS 2022 · 131 citations
