CLIPood: Generalizing CLIP to Out-of-Distributions
Yang Shu, Xingzhuo Guo, Jialong Wu, Ximei Wang, Jianmin Wang, Mingsheng Long
Abstract
Out-of-distribution (OOD) generalization, where the model needs to handle distribution shifts from training, is a major challenge of machine learning. Contrastive language-image pre-training (CLIP) models have shown impressive zero-shot ability, but the further adaptation of CLIP on downstream tasks undesirably degrades OOD performances. This paper aims at generalizing CLIP to out-of-distribution test data on downstream tasks. We propose CLIPood, a fine-tuning method that can adapt CLIP models to OOD situations where both domain shifts and open classes may occur on the unseen test data. To exploit the semantic relations between classes from the text modality, CLIPood introduces a new training objective, margin metric softmax (MMS), with class adaptive margins for fine-tuning. To incorporate both pre-trained zero-shot model and fine-tuned task-adaptive model, CLIPood leverages a new optimization strategy, Beta moving average (BMA), to maintain a temporal ensemble weighted by Beta distribution. Experiments on diverse datasets with different OOD scenarios show that CLIPood consistently outperforms existing generalization techniques.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1991cd61-540d-4f9b-afbb-a934118a3523Cited by top-tier papers38
- Towards Calibrated Robust Fine-Tuning of Vision-Language ModelsChangdae Oh, Hyesu Lim, Mijoo Kim, Dongyoon Han et al.NeurIPS 2024 · 49 citations
- CLIPCEIL: Domain Generalization through CLIP via Channel rEfinement and Image-text aLignmentXi Yu, Shinjae Yoo, Yuewei LinNeurIPS 2024 · 36 citations
- Vision-Language Models are Strong Noisy Label DetectorsTong Wei, Hao-Tian Li, Chun-Shu Li, Jiang-Xin Shi et al.NeurIPS 2024 · 26 citations
- Adapting to Distribution Shift by Visual Domain Prompt GenerationZhixiang Chi, Li Gu, Tao Zhong, Huan Liu et al.ICLR 2024 · 23 citations
- Overcoming the Pitfalls of Vision-Language Model Finetuning for OOD GeneralizationYuhang Zang, Hanlin Goh, Joshua Susskind, Chen HuangICLR 2024 · 17 citations
Builds on19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath et al.ICCV 2021 · 2,294 citations
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 2,258 citations
Related papers
- BATCLIP: Bimodal Online Test-Time Adaptation for CLIPSarthak Kumar Maharana, Baoming Zhang, Leonid Karlinsky, Rogério Feris et al.ICCV 2025 · 2 citations
- CLIPTTA: Robust Contrastive Vision-Language Test-Time AdaptationMarc Lafon, Gustavo Adolfo Vargas Hakim, Clément Rambour, Christian Desrosiers et al.NeurIPS 2025 · 5 citations
- Amend to Alignment: Decoupled Prompt Tuning for Mitigating Spurious Correlation in Vision-Language ModelsJie Zhang, Xiaosong Ma, Song Guo, Peng Li et al.ICML 2024 · 10 citations
- Non-Contrastive Learning Meets Language-Image Pre-TrainingJinghao Zhou, Li Dong, Zhe Gan, Lijuan Wang et al.CVPR 2023
- OT-CLIP: Understanding and Generalizing CLIP via Optimal TransportLiangliang Shi, Jack Fan, Junchi YanICML 2024 · 11 citations
