Statistical Consistency and Generalization of Contrastive Representation Learning
Yuanfan Li, Xiyuan Wei, Tianbao Yang, Yiming Ying
Abstract
Contrastive representation learning (CRL) underpins many modern foundation models. Despite recent theoretical progress, existing analyses suffer from several key limitations: (i) the statistical consistency of CRL remains poorly understood; (ii) available generalization bounds deteriorate as the number of negative samples increases, contradicting the empirical benefits of large negative sets; and (iii) the retrieval performance of CRL has received limited theoretical attention. In this paper, we develop a unified statistical learning theory for CRL. For downstream tasks, we evaluate retrieval quality using an AUC-type population criterion and show that the contrastive loss is statistically consistent with optimal ranking. We further establish a calibration-style inequality that quantitatively relates excess contrastive risk to excess retrieval suboptimality. For upstream training, we study both supervised and self-supervised contrastive objectives and derive generalization bounds of order and , respectively, where denotes the number of negative samples and the number of anchor points. These bounds not only explain the empirical advantages of large negative sets but also reveal an explicit trade-off between and . Extensive experiments on large-scale vision--language models corroborate our theoretical predictions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 55150e16-9ea8-4b82-b723-5f51e6b1808cBuilds on25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 2,360 citations
Related papers
- On the Surrogate Gap between Contrastive and Supervised LossesHan Bao, Yoshihiro Nagano, Kento NozawaICML 2022 · 27 citations
- On the Generalization of Multi-modal Contrastive LearningQi Zhang, Yifei Wang, Yisen WangICML 2023 · 34 citations
- X-Sample Contrastive Loss: Improving Contrastive Learning with Sample Similarity GraphsVlad Sobal, Mark Ibrahim, Randall Balestriero, Vivien Cabannes et al.ICLR 2025
- Debiased Contrastive LearningChing-Yao Chuang, Joshua Robinson, Yen-Chen Lin, Antonio Torralba et al.NeurIPS 2020 · 761 citations
- Understanding Negative Samples in Instance Discriminative Self-supervised Representation LearningKento Nozawa, Issei SatoNeurIPS 2021 · 56 citations
