Open-Vocabulary Customization from CLIP via Data-Free Knowledge Distillation
Yongxian Wei, Zixuan Hu, Li Shen, Zhenyi Wang, Chun Yuan, Dacheng Tao
Abstract
Vision-language models such as CLIP have demonstrated strong zero-shot performance, but their considerable size and inefficient inference limit customizable deployment for users. While knowledge distillation is a solution, it still requires the original data, which is not always available due to copyrights and privacy concerns. For many users seeking open-vocabulary customization, Data-Free Knowledge Distillation (DFKD) emerges as a promising direction. Upon rethinking DFKD, we find that existing methods fail on CLIP due to their heavy reliance on BatchNorm layers, which are unexpectedly unusable in CLIP. Based on our findings, we adopt image-text matching to achieve DFKD for CLIP, enabling customization based on arbitrary class texts. This involves (i) inversing a surrogate dataset from CLIP based on text prompts; and (ii) distilling a student model from CLIP using the surrogate dataset. Specifically, we introduce style dictionary diversification to enhance the diversity of synthetic images. To prevent uncontrollable semantics introduced by diversification, we propose a class consistency maintaining strategy to ensure the consistency of synthetic images. Based on synthetic images with various styles, we further propose meta knowledge distillation to train the student model with good generalization ability. Moreover, we introduce a simple yet effective method to enable customization based on few example images. Comprehensive experiments showcase the superiority of our approach across twelve customized tasks, achieving a 9.33% improvement compared to existing DFKD methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ff67527b-04c4-4646-b3cd-c8d5102c7707Cited by top-tier papers11
- Adaptive Defense against Harmful Fine-Tuning for Large Language Models via Bayesian Data SchedulerZixuan Hu, Li Shen, Zhenyi Wang, Yongxian Wei et al.NeurIPS 2025 · 16 citations
- OptMerge: Unifying Multimodal LLM Capabilities and Modalities via Model MergingYongxian Wei, Runxi Cheng, Weike Jin, Enneng Yang et al.ICLR 2026 · 10 citations
- SCOUT: Teaching Pre-trained Language Models to Enhance Reasoning via Flow Chain-of-ThoughtGuanghao Li, Wenhao Jiang, Mingfeng Chen, Yan Li et al.NeurIPS 2025 · 8 citations
- Model Inversion with Layer-Specific Modeling and Alignment for Data-Free Continual LearningRuilin Tong, Haodong Lu, Yuhang Liu, Dong GongNeurIPS 2025 · 6 citations
- Text-Guided Visual Prompt DINO for Generic SegmentationYuchen Guan, Chong Sun, Canmiao Fu, Zhipeng Huang et al.ICCV 2025 · 3 citations
Builds on45
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
Related papers
- Distilling Vision-Language Foundation Models: A Data-Free Approach via Prompt DiversificationYunyi Xuan, Weijie Chen, Shicai Yang, Di Xie et al.ACM MM 2023 · 4 citations
- Global Knowledge Calibration for Fast Open-Vocabulary SegmentationKunyang Han, Yong Liu, Jun Hao Liew, Henghui Ding et al.ICCV 2023 · 56 citations
- Multimodal Dataset Distillation Made Simple by Prototype-Guided Data SynthesisJunhyeok Choi, Sangwoo Mo, Minwoo ChaeICLR 2026
- Source-Free Domain Adaptation with Frozen Multimodal Foundation ModelSong Tang, Wenxin Su, Mao Ye, Xiatian ZhuCVPR 2024
- KAID: Knowledge-Aware Interactive Distillation for Vision-Language ModelsDa Zhang, Feiyu Wang, Bingyu Li, Zhiyuan Zhao et al.ACM MM 2025 · 10 citations
