Transferring Pre-trained Multimodal Representations with Cross-modal Similarity Matching
Byoungjip Kim, Sungik Choi, Dasol Hwang, Moontae Lee, Honglak Lee
摘要
Despite surprising performance on zero-shot transfer, pre-training a large-scale multimodal model is often prohibitive as it requires a huge amount of data and computing resources. In this paper, we propose a method (BeamCLIP) that can effectively transfer the representations of a large pre-trained multimodal model (CLIP-ViT) into a small target model (e.g., ResNet-18). For unsupervised transfer, we introduce cross-modal similarity matching (CSM) that enables a student model to learn the representations of a teacher model by matching the relative similarity distribution across text prompt embeddings. To better encode the text prompts, we design context-based prompt augmentation (CPA) that can alleviate the lexical ambiguity of input text prompts. Our experiments show that unsupervised representation transfer of a pre-trained vision-language model enables a small ResNet-18 to achieve a better ImageNet-1K top-1 linear probe accuracy (66.2%) than vision-only self-supervised learning (SSL) methods (e.g., SimCLR: 51.8%, SwAV: 63.7%), while closing the gap with supervised learning (69.8%). 36th Conference on Neural Information Processing Systems (NeurIPS 2022).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Factorized Contrastive Learning: Going Beyond Multi-view RedundancyPaul Pu Liang, Zihao Deng, Martin Q. Ma, James Y. Zou 等NeurIPS 2023 · 被引用 137 次
- DIME-FM : DIstilling Multimodal and Efficient Foundation ModelsXimeng Sun, Pengchuan Zhang, Peizhao Zhang, Hardik Shah 等ICCV 2023 · 被引用 42 次
- CLIP-CID: Efficient CLIP Distillation via Cluster-Instance DiscriminationKaicheng Yang, Tiancheng Gu, Xiang An, Haiqiang Jiang 等AAAI 2025 · 被引用 26 次
- Structural Information Guided Multimodal Pre-training for Vehicle-Centric PerceptionXiao Wang, Wentao Wu, Chenglong Li, Zhicheng Zhao 等AAAI 2024 · 被引用 10 次
- CAMILA: Context-Aware Masking for Image Editing with Language AlignmentHyunseung Kim, Chiho Choi, Srikanth Malla, Sai Prahladh Padmanabhan 等NeurIPS 2025 · 被引用 3 次
它引用的顶会 Paper19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
相关 Paper
- SimCLIP: Refining Image-Text Alignment with Simple Prompts for Zero-/Few-shot Anomaly DetectionChenghao Deng, Haote Xu, Xiaolu Chen, Haodi Xu 等ACM MM 2024 · 被引用 10 次
- Intra-Modal Proxy Learning for Zero-Shot Visual Categorization with CLIPQi Qian, Yuanhong Xu, Juhua HuNeurIPS 2023 · 被引用 34 次
- ViLT-CLIP: Video and Language Tuning CLIP with Multimodal Prompt Learning and Scenario-Guided OptimizationHao Wang, Fang Liu, Licheng Jiao, Jiahao Wang 等AAAI 2024 · 被引用 54 次
- RankCLIP: Ranking-Consistent Language-Image PretrainingYiming Zhang, Zhuokai Zhao, Zhaorun Chen, Zhili Feng 等ICCV 2025 · 被引用 1 次
- Language-Driven Cross-Modal Classifier for Zero-Shot Multi-Label Image RecognitionYicheng Liu, Jie Wen, Chengliang Liu, Xiaozhao Fang 等ICML 2024 · 被引用 7 次
