Leveraging Vision-Language Models for Improving Domain Generalization in Image Classification
Sravanti Addepalli, Ashish Ramayee Asokan, Lakshay Sharma, R. Venkatesh Babu
Abstract
Vision-Language Models (VLMs) such as CLIP are trained on large amounts of image-text pairs, resulting in remarkable generalization across several data distributions. However, in several cases, their expensive training and data collection/curation costs do not justify the end application. This motivates a vendor-client paradigm, where a vendor trains a large-scale VLM and grants only input-output access to clients on a pay-per-query basis in a black-box setting. The client aims to minimize inference cost by distilling the VLM to a student model using the limited available task-specific data, and further deploying this student model in the downstream application. While naive distillation largely improves the In-Domain (ID) accuracy of the student, it fails to transfer the superior outof-distribution (OOD) generalization of the VLM teacher using the limited available labeled images. To mitigate this, we propose Vision-Language to Vision -Align, Distill, Predict (VL2V-ADiP), which first aligns the vision and language modalities of the teacher model with the vision modality of a pre-trained student model, and further distills the aligned VLM representations to the student. This maximally retains the pre-trained features of the student, while also incorporating the rich representations of the VLM image encoder and the superior generalization of the text embeddings. The proposed approach achieves state-of-the-art results on the standard Domain Generalization benchmarks in a black-box teacher setting as well as a white-box setting where the weights of the VLM are accessible.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e8f6ac1e-b27d-41d0-8df9-53485e1ca692Cited by top-tier papers18
- CLIPCEIL: Domain Generalization through CLIP via Channel rEfinement and Image-text aLignmentXi Yu, Shinjae Yoo, Yuewei LinNeurIPS 2024 · 36 citations
- Reasoning-Driven Multimodal LLM for Domain GeneralizationZhipeng Xu, Zilong Wang, Xinyang Jiang, Dongsheng Li et al.ICLR 2026 · 11 citations
- Self-Refining Vision Language Model for Robotic Failure Detection and ReasoningCarl Qi, Xiaojie Wang, Silong Yong, Stephen Sheng et al.ICLR 2026 · 8 citations
- Learning a Cross-Modal Schrödinger Bridge for Visual Domain GeneralizationHao Zheng, Jingjun Yi, Qi Bi, Huimin Huang et al.NeurIPS 2025 · 1 citation
- Cross-Modal Knowledge Distillation without Paired Data: Theoretical Foundation and AlgorithmT. K Tran, Duc Chu Anh, Quang Hung Pham, Phi Le Nguyen et al.ICML 2026
Builds on39
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Moment Matching for Multi-Source Domain AdaptationXingchao Peng, Qinxun Bai, Xide Xia, Zijun Huang et al.ICCV 2019 · 2,239 citations
- In Search of Lost Domain GeneralizationIshaan Gulrajani, David Lopez-PazICLR 2021 · 1,416 citations
Related papers
- A Sentence Speaks a Thousand Images: Domain Generalization through Distilling CLIP with Language GuidanceZeyi Huang, Andy Zhou, Zijian Lin, Mu Cai et al.ICCV 2023 · 56 citations
- KAID: Knowledge-Aware Interactive Distillation for Vision-Language ModelsDa Zhang, Feiyu Wang, Bingyu Li, Zhiyuan Zhao et al.ACM MM 2025 · 10 citations
- Source-Free Domain Adaptation with Frozen Multimodal Foundation ModelSong Tang, Wenxin Su, Mao Ye, Xiatian ZhuCVPR 2024
- Domain Generalization in CLIP via Learning with Diverse Text PromptsChangsong Wen, Zelin Peng, Yu Huang, Xiaokang Yang et al.CVPR 2025
- PromptKD: Unsupervised Prompt Distillation for Vision-Language ModelsZheng Li, Xiang Li, Xinyi Fu, Xin Zhang et al.CVPR 2024
