KAID: Knowledge-Aware Interactive Distillation for Vision-Language Models
Da Zhang, Feiyu Wang, Bingyu Li, Zhiyuan Zhao, Junyu Gao, Xuelong Li
摘要
Vision-Language Models (VLMs) such as CLIP have demonstrated outstanding performance in cross-modal tasks, but the prohibitive computational cost hinders practical deployment. Although Knowledge Distillation (KD) provides a promising compression paradigm, most existing methods rely heavily on feature imitation and contrastive relations without explicit fine-grained alignment. Additionally, they do not fully leverage the multimodal interaction knowledge from the teacher model, restricting cross-modal semantic alignment. To address these challenges, we propose KAID, a Knowledge-Aware Interactive Distillation method for VLMs. Specifically, we first pretrain a large CLIP teacher model with domain few-shot labels and store text features as category vectors. Then, an Image Feature Matching (IFM) module is introduced to calculate the feature distribution of teacher-student models with improved cosine similarity, which achieves hierarchical knowledge transfer from global to local levels and enhances the fine-grained perception of student model through joint optimization. Moreover, a Pixel-Wise Alignment (PWA) module is constructed between the teacher's text features and the student's image features, employing a cross-modal attention mechanism to establish semantic associations, while a Text-guided Pixel alignment Loss function (TPloss) is concurrently designed to enhance the student's comprehension capabilities. Ultimately, the well-trained student model is used for inference. Extensive experiments on 11 datasets validate the effectiveness of our method. Specifically, our method achieves average improvements of 2.14% and 2.40% on the base and new classes across these datasets.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- CLIP-KD: An Empirical Study of CLIP Model DistillationChuanguang Yang, Zhulin An, Libo Huang, Junyu Bi 等CVPR 2024 · 被引用 50 次
- PromptKD: Unsupervised Prompt Distillation for Vision-Language ModelsZheng Li, Xiang Li, Xinyi Fu, Xin Zhang 等CVPR 2024
- CLIP-CID: Efficient CLIP Distillation via Cluster-Instance DiscriminationKaicheng Yang, Tiancheng Gu, Xiang An, Haiqiang Jiang 等AAAI 2025 · 被引用 26 次
- No Head Left Behind - Multi-Head Alignment Distillation for TransformersTianyang Zhao, Kunwar Yashraj Singh, Srikar Appalaraju, Peng Tang 等AAAI 2024 · 被引用 5 次
- HieRD: Hierarchical Relational Distillation for Vision-Language Embedding ModelsVinh Le, Nguyen Dang, Tu Vu, Linh Van 等ICML 2026
