CLIPCEIL: Domain Generalization through CLIP via Channel rEfinement and Image-text aLignment
Xi Yu, Shinjae Yoo, Yuewei Lin
摘要
Domain generalization (DG) is a fundamental yet challenging topic in machine learning. Recently, the remarkable zero-shot capabilities of the large pre-trained vision-language model (e.g., CLIP) have made it popular for various downstream tasks. However, the effectiveness of this capacity often degrades when there are shifts in data distribution during testing compared to the training data. In this paper, we propose a novel method, known as CLIPCEIL, a model that utilizes Channel rEfinement and Image-text aLignment to facilitate the CLIP to the inaccessible out-of-distribution test datasets that exhibit domain shifts. Specifically, we refine the feature channels in the visual domain to ensure they contain domain-invariant and class-relevant features by using a lightweight adapter. This is achieved by minimizing the inter-domain variance while maximizing the inter-class variance. In the meantime, we ensure the image-text alignment by aligning text embeddings of the class descriptions and their corresponding image embedding while further removing the domain-specific features. Moreover, our model integrates multi-scale CLIP features by utilizing a self-attention fusion module, technically implemented through one Transformer layer. Extensive experiments on five widely used benchmark datasets demonstrate that CLIPCEIL outperforms the existing state-of-the-art methods. The source code is available at https://github.com/yuxi120407/CLIPCEIL .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- How Do Optical Flow and Textual Prompts Collaborate to Assist in Audio-Visual Semantic Segmentation?Yujian Lee, Peng Gao, Yongqi Xu, Wentao FanICCV 2025 · 被引用 2 次
- Learning a Cross-Modal Schrödinger Bridge for Visual Domain GeneralizationHao Zheng, Jingjun Yi, Qi Bi, Huimin Huang 等NeurIPS 2025 · 被引用 1 次
- QT-DoG: Quantization-Aware Training for Domain GeneralizationSaqib Javed, Hieu Le, Mathieu SalzmannICML 2025
- Unlearning during Training: Domain-Specific Gradient Ascent for Domain GeneralizationDi Zhao, Jingfeng Zhang, Hongsheng Hu, Philippe Fournier-Viger 等ICLR 2026
- When and How Does CLIP Enable Domain and Compositional Generalization?Elias Kempf, Simon Schrodi, Max Argus, Thomas BroxICML 2025
它引用的顶会 Paper31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
相关 Paper
- Align Your Prompts: Test-Time Prompting with Distribution Alignment for Zero-Shot GeneralizationJameel Abdul Samadh, Hanan Gani, Noor Hussein, Muhammad Uzair Khattak 等NeurIPS 2023 · 被引用 147 次
- Domain Generalization in CLIP via Learning with Diverse Text PromptsChangsong Wen, Zelin Peng, Yu Huang, Xiaokang Yang 等CVPR 2025
- Efficient and Long-Tailed Generalization for Pre-trained Vision-Language ModelJiang-Xin Shi, Chi Zhang, Tong Wei, Yufeng LiKDD 2024 · 被引用 3 次
- A Sentence Speaks a Thousand Images: Domain Generalization through Distilling CLIP with Language GuidanceZeyi Huang, Andy Zhou, Zijian Lin, Mu Cai 等ICCV 2023 · 被引用 56 次
- UMFC: Unsupervised Multi-Domain Feature Calibration for Vision-Language ModelsJiachen Liang, Ruibing Hou, Minyang Hu, Hong Chang 等NeurIPS 2024 · 被引用 4 次
