TinyCLIP: CLIP Distillation via Affinity Mimicking and Weight Inheritance
Kan Wu, Houwen Peng, Zhenghong Zhou, Bin Xiao, Mengchen Liu, Lu Yuan, Hong Xuan, Michael Valenzuela, Xi Stephen Chen, Xinggang Wang, Hongyang Chao, Han Hu
摘要
In this paper, we propose a novel cross-modal distillation method, called TinyCLIP, for large-scale language-image pre-trained models. The method introduces two core techniques: affinity mimicking and weight inheritance. Affinity mimicking explores the interaction between modalities during distillation, enabling student models to mimic teachers' behavior of learning cross-modal feature alignment in a visual-linguistic affinity space. Weight inheritance transmits the pre-trained weights from the teacher models to their student counterparts to improve distillation efficiency. Moreover, we extend the method into a multi-stage progressive distillation to mitigate the loss of informative weights during extreme compression. Comprehensive experiments demonstrate the efficacy of TinyCLIP, showing that it can reduce the size of the pre-trained CLIP ViT-B/32 by 50%, while maintaining comparable zero-shot performance. While aiming for comparable performance, distillation with weight inheritance can speed up the training by 1.4 -7.8× compared to training from scratch. Moreover, our TinyCLIP ViT-8M/16, trained on YFCC-15M, achieves an impressive zero-shot top-1 accuracy of 41.1% on Im-ageNet, surpassing the original CLIP ViT-B/16 by 3.5% while utilizing only 8.9% parameters. Finally, we demonstrate the good transferability of TinyCLIP in various downstream tasks. Code and models will be open-sourced at aka.ms/tinyclip. * Equal contribution. Kan and Zhenghong were interns of Microsoft.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper33
- CLIP-KD: An Empirical Study of CLIP Model DistillationChuanguang Yang, Zhulin An, Libo Huang, Junyu Bi 等CVPR 2024 · 被引用 50 次
- MobileCLIP: Fast Image-Text Models through Multi-Modal Reinforced TrainingPavan Kumar Anasosalu Vasu, Hadi Pouransari, Fartash Faghri, Raviteja Vemulapalli 等CVPR 2024 · 被引用 29 次
- Implicit Multimodal Alignment: On the Generalization of Frozen LLMs to Multimodal InputsMustafa Shukor, Matthieu CordNeurIPS 2024 · 被引用 27 次
- CLIP-CID: Efficient CLIP Distillation via Cluster-Instance DiscriminationKaicheng Yang, Tiancheng Gu, Xiang An, Haiqiang Jiang 等AAAI 2025 · 被引用 26 次
- Generalized Contrastive Learning for Universal Multimodal RetrievalJungsoo Lee, Janghoon Cho, Hyojin Park, Durga Malladi 等NeurIPS 2025 · 被引用 11 次
它引用的顶会 Paper29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
相关 Paper
- CLIPPING: Distilling CLIP-Based Models with a Student Base for Video-Language RetrievalRenjing Pei, Jianzhuang Liu, Weimian Li, Bin Shao 等CVPR 2023
- KAID: Knowledge-Aware Interactive Distillation for Vision-Language ModelsDa Zhang, Feiyu Wang, Bingyu Li, Zhiyuan Zhao 等ACM MM 2025 · 被引用 10 次
- mCLIP: Multilingual CLIP via Cross-lingual TransferGuanhua Chen, Lu Hou, Yun Chen, Wenliang Dai 等ACL 2023 · 被引用 13 次
- MaskCLIP: Masked Self-Distillation Advances Contrastive Language-Image PretrainingXiaoyi Dong, Jianmin Bao, Yinglin Zheng, Ting Zhang 等CVPR 2023
- Building Vision-Language Models on Solid Foundations with Masked DistillationSepehr Sameni, Kushal Kafle, Hao Tan, Simon JenniCVPR 2024 · 被引用 4 次
