CLIPCEIL: Domain Generalization through CLIP via Channel rEfinement and Image-text aLignment
Xi Yu, Shinjae Yoo, Yuewei Lin
Abstract
Domain generalization (DG) is a fundamental yet challenging topic in machine learning. Recently, the remarkable zero-shot capabilities of the large pre-trained vision-language model (e.g., CLIP) have made it popular for various downstream tasks. However, the effectiveness of this capacity often degrades when there are shifts in data distribution during testing compared to the training data. In this paper, we propose a novel method, known as CLIPCEIL, a model that utilizes Channel rEfinement and Image-text aLignment to facilitate the CLIP to the inaccessible out-of-distribution test datasets that exhibit domain shifts. Specifically, we refine the feature channels in the visual domain to ensure they contain domain-invariant and class-relevant features by using a lightweight adapter. This is achieved by minimizing the inter-domain variance while maximizing the inter-class variance. In the meantime, we ensure the image-text alignment by aligning text embeddings of the class descriptions and their corresponding image embedding while further removing the domain-specific features. Moreover, our model integrates multi-scale CLIP features by utilizing a self-attention fusion module, technically implemented through one Transformer layer. Extensive experiments on five widely used benchmark datasets demonstrate that CLIPCEIL outperforms the existing state-of-the-art methods. The source code is available at https://github.com/yuxi120407/CLIPCEIL .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 686c4572-7e92-4495-a2ea-d0e7c67706e0Cited by top-tier papers6
- How Do Optical Flow and Textual Prompts Collaborate to Assist in Audio-Visual Semantic Segmentation?Yujian Lee, Peng Gao, Yongqi Xu, Wentao FanICCV 2025 · 2 citations
- Learning a Cross-Modal Schrödinger Bridge for Visual Domain GeneralizationHao Zheng, Jingjun Yi, Qi Bi, Huimin Huang et al.NeurIPS 2025 · 1 citation
- QT-DoG: Quantization-Aware Training for Domain GeneralizationSaqib Javed, Hieu Le, Mathieu SalzmannICML 2025
- Unlearning during Training: Domain-Specific Gradient Ascent for Domain GeneralizationDi Zhao, Jingfeng Zhang, Hongsheng Hu, Philippe Fournier-Viger et al.ICLR 2026
- When and How Does CLIP Enable Domain and Compositional Generalization?Elias Kempf, Simon Schrodi, Max Argus, Thomas BroxICML 2025
Builds on31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
Related papers
- Align Your Prompts: Test-Time Prompting with Distribution Alignment for Zero-Shot GeneralizationJameel Abdul Samadh, Hanan Gani, Noor Hussein, Muhammad Uzair Khattak et al.NeurIPS 2023 · 147 citations
- Domain Generalization in CLIP via Learning with Diverse Text PromptsChangsong Wen, Zelin Peng, Yu Huang, Xiaokang Yang et al.CVPR 2025
- Efficient and Long-Tailed Generalization for Pre-trained Vision-Language ModelJiang-Xin Shi, Chi Zhang, Tong Wei, Yufeng LiKDD 2024 · 3 citations
- A Sentence Speaks a Thousand Images: Domain Generalization through Distilling CLIP with Language GuidanceZeyi Huang, Andy Zhou, Zijian Lin, Mu Cai et al.ICCV 2023 · 56 citations
- UMFC: Unsupervised Multi-Domain Feature Calibration for Vision-Language ModelsJiachen Liang, Ruibing Hou, Minyang Hu, Hong Chang et al.NeurIPS 2024 · 4 citations
