A Sentence Speaks a Thousand Images: Domain Generalization through Distilling CLIP with Language Guidance
Zeyi Huang, Andy Zhou, Zijian Lin, Mu Cai, Haohan Wang, Yong Jae Lee
Abstract
Domain generalization studies the problem of training a model with samples from several domains (or distributions) and then testing the model with samples from a new, unseen domain. In this paper, we propose a novel approach for domain generalization that leverages recent advances in large vision-language models, specifically a CLIP teacher model, to train a smaller model that generalizes to unseen domains. The key technical contribution is a new type of regularization that requires the student’s learned image representations to be close to the teacher’s learned text representations obtained from encoding the corresponding text descriptions of images. We introduce two designs of the loss function, absolute and relative distance, which provide specific guidance on how the training process of the student model should be regularized. We evaluate our proposed method, dubbed RISE (Regularized Invariance with Semantic Embeddings), on various benchmark datasets, and show that it outperforms several state-of-the-art domain generalization methods. To our knowledge, our work is the first to leverage knowledge distillation using a large vision-language model for domain generalization. By incorporating text-based information, RISE improves the generalization capability of machine learning models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fccdcc0a-c446-4ea1-9a3f-a903bad02535Cited by top-tier papers20
- Distilling Out-of-Distribution Robustness from Vision-Language Foundation ModelsAndy Zhou, Jindong Wang, Yu-Xiong Wang, Haohan WangNeurIPS 2023 · 14 citations
- Tell, Don't Show: Language Guidance Eases Transfer Across Domains in Images and VideosTarun Kalluri, Bodhisattwa Prasad Majumder, Manmohan ChandrakerICML 2024 · 7 citations
- UMFC: Unsupervised Multi-Domain Feature Calibration for Vision-Language ModelsJiachen Liang, Ruibing Hou, Minyang Hu, Hong Chang et al.NeurIPS 2024 · 4 citations
- Learning a Cross-Modal Schrödinger Bridge for Visual Domain GeneralizationHao Zheng, Jingjun Yi, Qi Bi, Huimin Huang et al.NeurIPS 2025 · 1 citation
- Customizing Domain Adapters for Domain GeneralizationYuyang Ji, Zeyi Huang, Haohan Wang, Yong Jae LeeICCV 2025 · 1 citation
Builds on26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 1,274 citations
- FILIP: Fine-grained Interactive Language-Image Pre-TrainingLewei Yao, Runhui Huang, Lu Hou, Guansong Lu et al.ICLR 2022 · 827 citations
- SWAD: Domain Generalization by Seeking Flat MinimaJunbum Cha, Sanghyuk Chun, Kyungjae Lee, Han-Cheol Cho et al.NeurIPS 2021 · 630 citations
Related papers
- Domain Generalization in CLIP via Learning with Diverse Text PromptsChangsong Wen, Zelin Peng, Yu Huang, Xiaokang Yang et al.CVPR 2025
- CLIPCEIL: Domain Generalization through CLIP via Channel rEfinement and Image-text aLignmentXi Yu, Shinjae Yoo, Yuewei LinNeurIPS 2024 · 36 citations
- Leveraging Vision-Language Models for Improving Domain Generalization in Image ClassificationSravanti Addepalli, Ashish Ramayee Asokan, Lakshay Sharma, R. Venkatesh BabuCVPR 2024
- PromptKD: Unsupervised Prompt Distillation for Vision-Language ModelsZheng Li, Xiang Li, Xinyi Fu, Xin Zhang et al.CVPR 2024
- Disentangled Prompt Representation for Domain GeneralizationDe Cheng, Zhipeng Xu, Xinyang Jiang, Nannan Wang et al.CVPR 2024
