When and How Does CLIP Enable Domain and Compositional Generalization?
Elias Kempf, Simon Schrodi, Max Argus, Thomas Brox
摘要
The remarkable generalization performance of contrastive vision-language models like CLIP is often attributed to the diversity of their training distributions. However, key questions remain unanswered: Can CLIP generalize to an entirely unseen domain when trained on a diverse mixture of domains (domain generalization)? Can it generalize to unseen classes within partially seen domains (compositional generalization)? What factors affect such generalization? To answer these questions, we trained CLIP models on systematically constructed training distributions with controlled domain diversity and object class exposure. Our experiments show that domain diversity is essential for both domain and compositional generalization, yet compositional generalization can be surprisingly weaker than domain generalization when the training distribution contains a suboptimal subset of the test domain. Through data-centric and mechanistic analyses, we find that successful generalization requires the learning of sufficiently shared representations in intermediate layers and circuits.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Bridging Domains through Subspace-Aware Model MergingLevy Chaves, Chao Zhou, Rebekka Burkholz, Eduardo Valle 等CVPR 2026 · 被引用 2 次
- Necessary Conditions for Compositional Generalization of Embedding ModelsArnas Uselis, Andrea Dittadi, Seong Joon OhICML 2026
- Are Object-Centric Representations Better at Compositional Generalization?Ferdinand Kapl, Amir Mohammad Karimi Mamaghan, Maximilian Seitzer, Karl Johansson 等ICML 2026
- How can embedding models bind concepts?Arnas Uselis, Darina Koishigarina, Seong Joon OhICML 2026
它引用的顶会 Paper32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
相关 Paper
- Does Data Scaling Lead to Visual Compositional Generalization?Arnas Uselis, Andrea Dittadi, Seong Joon OhICML 2025
- Left–Right Symmetry Breaking in CLIP-Style Vision-Language Models Trained on Synthetic Spatial-Relation DataTakaki Yamamoto, Chihiro Noguchi, Toshihiro TanizawaICML 2026
- Domain Generalization in CLIP via Learning with Diverse Text PromptsChangsong Wen, Zelin Peng, Yu Huang, Xiaokang Yang 等CVPR 2025
- What Makes CLIP More Robust to Long-Tailed Pre-Training Data? A Controlled Study for Transferable InsightsXin Wen, Bingchen Zhao, Yilun Chen, Jiangmiao Pang 等NeurIPS 2024 · 被引用 19 次
- Understanding the Robustness of Multi-modal Contrastive Learning to Distribution ShiftYihao Xue, Siddharth Joshi, Dang Nguyen, Baharan MirzasoleimanICLR 2024 · 被引用 6 次
