Why do We Need Large Batchsizes in Contrastive Learning? A Gradient-Bias Perspective
Changyou Chen, Jianyi Zhang, Yi Xu, Liqun Chen, Jiali Duan, Yiran Chen, Son Tran, Belinda Zeng, Trishul Chilimbi
Abstract
Contrastive learning (CL) has been the de facto technique for self-supervised representation learning (SSL), with impressive empirical success such as multi-modal representation learning. However, traditional CL loss only considers negative samples from a minibatch, which could cause biased gradients due to the nondecomposibility of the loss. For the first time, we consider optimizing a more generalized contrastive loss, where each data sample is associated with an infinite number of negative samples. We show that directly using minibatch stochastic optimization could lead to gradient bias. To remedy this, we propose an efficient Bayesian data augmentation technique to augment the contrastive loss into a decomposable one, where standard stochastic optimization can be directly applied without gradient bias. Specifically, our augmented loss defines a joint distribution over the model parameters and the augmented parameters, which can be conveniently optimized by a proposed stochastic expectation-maximization algorithm. Our framework is more general and is related to several popular SSL algorithms. We verify our framework on both small scale models and several large foundation models, including SSL of ImageNet and SSL for vision-language representation learning. Experiment results indicate the existence of gradient bias in all cases, and demonstrate the effectiveness of the proposed method on improving previous state of the arts. Remarkably, our method can outperform the strong MoCo-v3 under the same hyper-parameter setting with only around half of the minibatch size; and also obtains strong results in the recent public benchmark ELEVATER for few-shot image classification. * We will use auxiliary (random) variable and auxiliary data interchangeably without distinction, as they are equivalently used in the Bayesian data augmentation community.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 23ecd666-8f25-4ca0-baee-9dab4d12dafbCited by top-tier papers14
- 1000 Layer Networks for Self-Supervised RL: Scaling Depth Can Enable New Goal-Reaching CapabilitiesKevin Wang, Ishaan Javali, Michal Bortkiewicz, Tomasz Trzcinski et al.NeurIPS 2025 · 46 citations
- Middle-Layer Representation Alignment for Cross-Lingual Transfer in Fine-Tuned LLMsDanni Liu, Jan NiehuesACL 2025 · 23 citations
- J4R: Learning to Judge with Equivalent Initial State Group Relative Policy OptimizationAustin Xu, Yilun Zhou, Xuan-Phi Nguyen, Caiming Xiong et al.ACL 2026 · 8 citations
- Hierarchical Pretraining on Multimodal Electronic Health RecordsXiaochen Wang, Junyu Luo, Jiaqi Wang, Ziyi Yin et al.EMNLP 2023 · 7 citations
- Bridging Mini-Batch and Asymptotic Analysis in Contrastive Learning: From InfoNCE to Kernel-Based LossesPanagiotis Koromilas, Giorgos Bouritsas, Theodoros Giannakopoulos, Mihalis Nicolaou et al.ICML 2024 · 7 citations
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
Related papers
- Maximizing Incremental Information Entropy for Contrastive LearningJiansong Zhang, Zhuoqin Yang, Xu Wu, Xiaoling Luo et al.ICLR 2026
- AUC-CL: A Batchsize-Robust Framework for Self-Supervised Contrastive Representation LearningRohan Sharma, Kaiyi Ji, Zhiqiang Xu, Changyou ChenICLR 2024 · 5 citations
- Contrastive Learning with Adversarial ExamplesChih-Hui Ho, Nuno VasconcelosNeurIPS 2020 · 174 citations
- Variational Supervised Contrastive LearningZiwen Wang, Jiajun Fan, Thao Nguyen, Heng Ji et al.NeurIPS 2025 · 7 citations
- Modal-aware Bias Constrained Contrastive Learning for Multimodal RecommendationWei Yang, Zhengru Fang, Tianle Zhang, Shiguang Wu et al.ACM MM 2023 · 15 citations
