On the Generalization of Multi-modal Contrastive Learning
Qi Zhang, Yifei Wang, Yisen Wang
摘要
Multi-modal contrastive learning (MMCL) has recently garnered considerable interest due to its superior performance in visual tasks, achieved by embedding multi-modal data, such as visual-language pairs. However, there still lack theoretical understandings of how MMCL extracts useful visual representation from multi-modal pairs, and particularly, how MMCL outperforms previous approaches like self-supervised contrastive learning (SSCL). In this paper, by drawing an intrinsic connection between MMCL and asymmetric matrix factorization, we establish the first generalization guarantees of MMCL for visual downstream tasks. Based on this framework, we further unify MMCL and SSCL by showing that MMCL implicitly performs SSCL with (pseudo) positive pairs induced by text pairs. Through this unified perspective, we characterize the advantage of MMCL by showing that text pairs induce more semantically consistent and diverse positive pairs, which, according to our analysis, provably benefit downstream generalization. Inspired by this finding, we propose CLIP-guided resampling methods to significantly improve the downstream performance of SSCL on ImageNet by leveraging multi-modal information. Code is available at https://github.com/PKU-ML/CLIP-Help-SimCLR.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper23
- Towards Calibrated Robust Fine-Tuning of Vision-Language ModelsChangdae Oh, Hyesu Lim, Mijoo Kim, Dongyoon Han 等NeurIPS 2024 · 被引用 49 次
- A Sober Look at the Robustness of CLIPs to Spurious FeaturesQizhou Wang, Yong Lin, Yongqiang Chen, Ludwig Schmidt 等NeurIPS 2024 · 被引用 46 次
- Do Generated Data Always Help Contrastive Learning?Yifei Wang, Jizhe Zhang, Yisen WangICLR 2024 · 被引用 36 次
- Continual Multimodal Contrastive LearningXiaohao Liu, Xiaobo Xia, See-Kiong Ng, Tat-Seng ChuaNeurIPS 2025 · 被引用 25 次
- Adversarial Examples Are Not Real FeaturesAng Li, Yifei Wang, Yiwen Guo, Yisen WangNeurIPS 2023 · 被引用 24 次
它引用的顶会 Paper16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 被引用 2,360 次
相关 Paper
- Understanding Transferable Representation Learning and Zero-shot Transfer in CLIPZixiang Chen, Yihe Deng, Yuanzhi Li, Quanquan GuICLR 2024 · 被引用 21 次
- Generalized Contrastive Learning for Universal Multimodal RetrievalJungsoo Lee, Janghoon Cho, Hyojin Park, Durga Malladi 等NeurIPS 2025 · 被引用 11 次
- Multi-modal contrastive learning adapts to intrinsic dimensions of shared latent variablesYu Gui, Cong Ma, Zongming MaNeurIPS 2025 · 被引用 9 次
- RankCLIP: Ranking-Consistent Language-Image PretrainingYiming Zhang, Zhuokai Zhao, Zhaorun Chen, Zhili Feng 等ICCV 2025 · 被引用 1 次
- Statistical Consistency and Generalization of Contrastive Representation LearningYuanfan Li, Xiyuan Wei, Tianbao Yang, Yiming YingICML 2026 · 被引用 1 次
