Does CLIP's generalization performance mainly stem from high train-test similarity?
Prasanna Mayilvahanan, Thaddäus Wiedemer, Evgenia Rusak, Matthias Bethge, Wieland Brendel
摘要
Foundation models like CLIP are trained on hundreds of millions of samples and effortlessly generalize to new tasks and inputs. Out of the box, CLIP shows stellar zero-shot and few-shot capabilities on a wide range of out-of-distribution (OOD) benchmarks, which prior works attribute mainly to today's large and comprehensive training dataset (like LAION). However, it is questionable how meaningful terms like out-of-distribution generalization are for CLIP as it seems likely that web-scale datasets like LAION simply contain many samples that are similar to common OOD benchmarks originally designed for ImageNet. To test this hypothesis, we retrain CLIP on pruned LAION splits that replicate ImageNet's train-test similarity with respect to common OOD benchmarks. While we observe a performance drop on some benchmarks, surprisingly, CLIP's overall performance remains high. This shows that high train-test similarity is insufficient to explain CLIP's OOD performance, and other properties of the training data must drive CLIP to learn more generalizable representations. Additionally, by pruning data points that are dissimilar to the OOD benchmarks, we uncover a 100M split of LAION (th of its original size) on which CLIP can be trained to match its original OOD performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper23
- No "Zero-Shot" Without Exponential Data: Pretraining Concept Frequency Determines Multimodal Model PerformanceVishaal Udandarao, Ameya Prabhu, Adhiraj Ghosh, Yash Sharma 等NeurIPS 2024 · 被引用 101 次
- Frustratingly Easy Test-Time Adaptation of Vision-Language ModelsMatteo Farina, Gianni Franchi, Giovanni Iacca, Massimiliano Mancini 等NeurIPS 2024 · 被引用 47 次
- A Sober Look at the Robustness of CLIPs to Spurious FeaturesQizhou Wang, Yong Lin, Yongqiang Chen, Ludwig Schmidt 等NeurIPS 2024 · 被引用 46 次
- On the Comparison between Multi-modal and Single-modal Contrastive LearningWei Huang, Andi Han, Yongqiang Chen, Yuan Cao 等NeurIPS 2024 · 被引用 26 次
- What Makes CLIP More Robust to Long-Tailed Pre-Training Data? A Controlled Study for Transferable InsightsXin Wen, Bingchen Zhao, Yilun Chen, Jiangmiao Pang 等NeurIPS 2024 · 被引用 19 次
它引用的顶会 Paper9
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath 等ICCV 2021 · 被引用 2,294 次
- Data Determines Distributional Robustness in Contrastive Language Image Pre-training (CLIP)Alex Fang, Gabriel Ilharco, Mitchell Wortsman, Yuhao Wan 等ICML 2022 · 被引用 183 次
相关 Paper
- In Search of Forgotten Domain GeneralizationPrasanna Mayilvahanan, Roland S. Zimmermann, Thaddäus Wiedemer, Evgenia Rusak 等ICLR 2025
- Effective pruning of web-scale datasets based on complexity of concept clustersAmro Abbas, Evgenia Rusak, Kushal Tirumala, Wieland Brendel 等ICLR 2024 · 被引用 30 次
- Reproducible Scaling Laws for Contrastive Language-Image LearningMehdi Cherti, Romain Beaumont, Ross Wightman, Mitchell Wortsman 等CVPR 2023
- Quality Not Quantity: On the Interaction between Dataset Design and Robustness of CLIPThao Nguyen, Gabriel Ilharco, Mitchell Wortsman, Sewoong Oh 等NeurIPS 2022 · 被引用 131 次
- Reproducible Vision-Language Models Meet Concepts Out of Pre-TrainingZiliang Chen, Xin Huang, Xiaoxuan Fan, Keze Wang 等CVPR 2025
