Does CLIP's generalization performance mainly stem from high train-test similarity?
Prasanna Mayilvahanan, Thaddäus Wiedemer, Evgenia Rusak, Matthias Bethge, Wieland Brendel
Abstract
Foundation models like CLIP are trained on hundreds of millions of samples and effortlessly generalize to new tasks and inputs. Out of the box, CLIP shows stellar zero-shot and few-shot capabilities on a wide range of out-of-distribution (OOD) benchmarks, which prior works attribute mainly to today's large and comprehensive training dataset (like LAION). However, it is questionable how meaningful terms like out-of-distribution generalization are for CLIP as it seems likely that web-scale datasets like LAION simply contain many samples that are similar to common OOD benchmarks originally designed for ImageNet. To test this hypothesis, we retrain CLIP on pruned LAION splits that replicate ImageNet's train-test similarity with respect to common OOD benchmarks. While we observe a performance drop on some benchmarks, surprisingly, CLIP's overall performance remains high. This shows that high train-test similarity is insufficient to explain CLIP's OOD performance, and other properties of the training data must drive CLIP to learn more generalizable representations. Additionally, by pruning data points that are dissimilar to the OOD benchmarks, we uncover a 100M split of LAION (th of its original size) on which CLIP can be trained to match its original OOD performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 766b598f-480f-408c-91ef-c9ff058706dbCited by top-tier papers23
- No "Zero-Shot" Without Exponential Data: Pretraining Concept Frequency Determines Multimodal Model PerformanceVishaal Udandarao, Ameya Prabhu, Adhiraj Ghosh, Yash Sharma et al.NeurIPS 2024 · 101 citations
- Frustratingly Easy Test-Time Adaptation of Vision-Language ModelsMatteo Farina, Gianni Franchi, Giovanni Iacca, Massimiliano Mancini et al.NeurIPS 2024 · 47 citations
- A Sober Look at the Robustness of CLIPs to Spurious FeaturesQizhou Wang, Yong Lin, Yongqiang Chen, Ludwig Schmidt et al.NeurIPS 2024 · 46 citations
- On the Comparison between Multi-modal and Single-modal Contrastive LearningWei Huang, Andi Han, Yongqiang Chen, Yuan Cao et al.NeurIPS 2024 · 26 citations
- What Makes CLIP More Robust to Long-Tailed Pre-Training Data? A Controlled Study for Transferable InsightsXin Wen, Bingchen Zhao, Yilun Chen, Jiangmiao Pang et al.NeurIPS 2024 · 19 citations
Builds on9
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath et al.ICCV 2021 · 2,294 citations
- Data Determines Distributional Robustness in Contrastive Language Image Pre-training (CLIP)Alex Fang, Gabriel Ilharco, Mitchell Wortsman, Yuhao Wan et al.ICML 2022 · 183 citations
Related papers
- In Search of Forgotten Domain GeneralizationPrasanna Mayilvahanan, Roland S. Zimmermann, Thaddäus Wiedemer, Evgenia Rusak et al.ICLR 2025
- Effective pruning of web-scale datasets based on complexity of concept clustersAmro Abbas, Evgenia Rusak, Kushal Tirumala, Wieland Brendel et al.ICLR 2024 · 30 citations
- Reproducible Scaling Laws for Contrastive Language-Image LearningMehdi Cherti, Romain Beaumont, Ross Wightman, Mitchell Wortsman et al.CVPR 2023
- Quality Not Quantity: On the Interaction between Dataset Design and Robustness of CLIPThao Nguyen, Gabriel Ilharco, Mitchell Wortsman, Sewoong Oh et al.NeurIPS 2022 · 131 citations
- Reproducible Vision-Language Models Meet Concepts Out of Pre-TrainingZiliang Chen, Xin Huang, Xiaoxuan Fan, Keze Wang et al.CVPR 2025
