Building Vision-Language Models on Solid Foundations with Masked Distillation
Sepehr Sameni, Kushal Kafle, Hao Tan, Simon Jenni
摘要
Recent advancements in Vision-Language Models (VLMs) have marked a significant leap in bridging the gap between computer vision and natural language processing. However, traditional VLMs, trained through contrastive learning on limited and noisy image-text pairs, often lack the spatial and linguistic understanding to generalize well to dense vision tasks or less common languages. Our approach, Solid Foun-dation CLIP (SF-CLIP), circumvents this issue by implicitly building on the solid visual and language understanding of foundational models trained on vast amounts of unimodal data. SF-CLIP integrates contrastive image-text pretraining with a masked knowledge distillation from large foundational text and vision models. This methodology guides our VLM in developing robust text and image representations. As a result, SF-CLIP shows exceptional zero-shot classification accuracy and enhanced image and text retrieval capabilities, setting a new state of the art for ViT-B/16 trained on YFCC15M and CC12M.Moreover, the dense per-patch supervision enhances our zero-shot and linear probe performance in semantic segmentation tasks. A remarkable aspect of our model is its multilingual proficiency, evidenced by strong retrieval results in multiple languages despite being trained predominantly on English data. We achieve all of these improvements without sacrificing the training efficiency through our selective application of masked distillation and the inheritance of teacher word embeddings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Enhancing Compositional Reasoning in CLIP via Reconstruction and Alignment of Text DescriptionsJihoon Kwon, Kyle Min, Jy-yong SohnNeurIPS 2025 · 被引用 4 次
- Locality Alignment Improves Vision-Language ModelsIan Connick Covert, Tony Sun, James Zou, Tatsunori HashimotoICLR 2025
- Progressive Mask Distillation for Self-supervised Video RepresentationKewei Wu, Chong Liang, Zhao Xie, Dan GuoCVPR 2026
- COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-trainingSanghwan Kim, Rui Xiao, Mariana-Iuliana Georgescu, Stephan Alaniz 等CVPR 2025
- Active Data Curation Effectively Distills Large-Scale Multimodal ModelsVishaal Udandarao, Nikhil Parthasarathy, Muhammad Ferjad Naeem, Talfan Evans 等CVPR 2025
它引用的顶会 Paper26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
相关 Paper
- MaskCLIP: Masked Self-Distillation Advances Contrastive Language-Image PretrainingXiaoyi Dong, Jianmin Bao, Yinglin Zheng, Ting Zhang 等CVPR 2023
- mCLIP: Multilingual CLIP via Cross-lingual TransferGuanhua Chen, Lu Hou, Yun Chen, Wenliang Dai 等ACL 2023 · 被引用 13 次
- CLIP-KD: An Empirical Study of CLIP Model DistillationChuanguang Yang, Zhulin An, Libo Huang, Junyu Bi 等CVPR 2024 · 被引用 50 次
- How Much Can CLIP Benefit Vision-and-Language Tasks?Sheng Shen, Liunian Harold Li, Hao Tan, Mohit Bansal 等ICLR 2022 · 被引用 503 次
- RankCLIP: Ranking-Consistent Language-Image PretrainingYiming Zhang, Zhuokai Zhao, Zhaorun Chen, Zhili Feng 等ICCV 2025 · 被引用 1 次
