Distilling Large Vision-Language Model with Out-of-Distribution Generalizability
Xuanlin Li, Yunhao Fang, Minghua Liu, Zhan Ling, Zhuowen Tu, Hao Su
Abstract
Large vision-language models have achieved outstanding performance, but their size and computational requirements make their deployment on resource-constrained devices and time-sensitive tasks impractical. Model distillation, the process of creating smaller, faster models that maintain the performance of larger models, is a promising direction towards the solution. This paper investigates the distillation of visual representations in large teacher vision-language models into lightweight student models using a small-or mid-scale dataset. Notably, this study focuses on open-vocabulary outof-distribution (OOD) generalization, a challenging problem that has been overlooked in previous model distillation literature. We propose two principles from vision and language modality perspectives to enhance student's OOD generalization: (1) by better imitating teacher's visual representation space, and carefully promoting better coherence in vision-language alignment with the teacher; (2) by enriching the teacher's language representations with informative and finegrained semantic attributes to effectively distinguish between different labels. We propose several metrics and conduct extensive experiments to investigate their techniques. The results demonstrate significant improvements in zeroshot and few-shot student performance on open-vocabulary out-of-distribution classification, highlighting the effectiveness of our proposed approaches. Project poster: this link. Code: this link.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ef71ec4d-bec3-4c8f-aade-77dec73ea384Cited by top-tier papers15
- CLIP-KD: An Empirical Study of CLIP Model DistillationChuanguang Yang, Zhulin An, Libo Huang, Junyu Bi et al.CVPR 2024 · 50 citations
- Overcoming the Pitfalls of Vision-Language Model Finetuning for OOD GeneralizationYuhang Zang, Hanlin Goh, Joshua Susskind, Chen HuangICLR 2024 · 17 citations
- Knowledge Transfer from Vision Foundation Models for Efficient Training of Small Task-specific ModelsRaviteja Vemulapalli, Hadi Pouransari, Fartash Faghri, Sachin Mehta et al.ICML 2024 · 15 citations
- ImageNet-D: Benchmarking Neural Network Robustness on Diffusion Synthetic ObjectChenshuang Zhang, Fei Pan, Junmo Kim, In So Kweon et al.CVPR 2024 · 9 citations
- From Semantics, Scene to Instance-awareness: Distilling Foundation Model for Open-vocabulary Grounded Situation RecognitionChen Cai, Tianyi Liu, Jianjun Gao, Wenyang Liu et al.ACM MM 2025 · 2 citations
Builds on31
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
Related papers
- LLaVA-KD: A Framework of Distilling Multimodal Large Language ModelsYuxuan Cai, Jiangning Zhang, Haoyang He, Xinwei He et al.ICCV 2025 · 9 citations
- Distilling DETR with Visual-Linguistic Knowledge for Open-Vocabulary Object DetectionLiangqi Li, Jiaxu Miao, Dahu Shi, Wenming Tan et al.ICCV 2023 · 35 citations
- Open-Vocabulary One-Stage Detection with Hierarchical Visual-Language Knowledge DistillationZongyang Ma, Guan Luo, Jin Gao, Liang Li et al.CVPR 2022 · 44 citations
- Dataset Distillation Via Vision-Language Category PrototypeYawen Zou, Guang Li, Duo Su, Zi Wang et al.ICCV 2025 · 2 citations
- ATAS: Any-to-Any Self-Distillation for Enhanced Open-Vocabulary Dense PredictionJuan Yeo, Soonwoo Cha, Jiwoo Song, Hyunbin Jin et al.ICCV 2025
