Why do We Need Large Batchsizes in Contrastive Learning? A Gradient-Bias Perspective
Changyou Chen, Jianyi Zhang, Yi Xu, Liqun Chen, Jiali Duan, Yiran Chen, Son Tran, Belinda Zeng, Trishul Chilimbi
摘要
Contrastive learning (CL) has been the de facto technique for self-supervised representation learning (SSL), with impressive empirical success such as multi-modal representation learning. However, traditional CL loss only considers negative samples from a minibatch, which could cause biased gradients due to the nondecomposibility of the loss. For the first time, we consider optimizing a more generalized contrastive loss, where each data sample is associated with an infinite number of negative samples. We show that directly using minibatch stochastic optimization could lead to gradient bias. To remedy this, we propose an efficient Bayesian data augmentation technique to augment the contrastive loss into a decomposable one, where standard stochastic optimization can be directly applied without gradient bias. Specifically, our augmented loss defines a joint distribution over the model parameters and the augmented parameters, which can be conveniently optimized by a proposed stochastic expectation-maximization algorithm. Our framework is more general and is related to several popular SSL algorithms. We verify our framework on both small scale models and several large foundation models, including SSL of ImageNet and SSL for vision-language representation learning. Experiment results indicate the existence of gradient bias in all cases, and demonstrate the effectiveness of the proposed method on improving previous state of the arts. Remarkably, our method can outperform the strong MoCo-v3 under the same hyper-parameter setting with only around half of the minibatch size; and also obtains strong results in the recent public benchmark ELEVATER for few-shot image classification. * We will use auxiliary (random) variable and auxiliary data interchangeably without distinction, as they are equivalently used in the Bayesian data augmentation community.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- 1000 Layer Networks for Self-Supervised RL: Scaling Depth Can Enable New Goal-Reaching CapabilitiesKevin Wang, Ishaan Javali, Michal Bortkiewicz, Tomasz Trzcinski 等NeurIPS 2025 · 被引用 46 次
- Middle-Layer Representation Alignment for Cross-Lingual Transfer in Fine-Tuned LLMsDanni Liu, Jan NiehuesACL 2025 · 被引用 23 次
- J4R: Learning to Judge with Equivalent Initial State Group Relative Policy OptimizationAustin Xu, Yilun Zhou, Xuan-Phi Nguyen, Caiming Xiong 等ACL 2026 · 被引用 8 次
- Hierarchical Pretraining on Multimodal Electronic Health RecordsXiaochen Wang, Junyu Luo, Jiaqi Wang, Ziyi Yin 等EMNLP 2023 · 被引用 7 次
- Bridging Mini-Batch and Asymptotic Analysis in Contrastive Learning: From InfoNCE to Kernel-Based LossesPanagiotis Koromilas, Giorgos Bouritsas, Theodoros Giannakopoulos, Mihalis Nicolaou 等ICML 2024 · 被引用 7 次
它引用的顶会 Paper23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
相关 Paper
- Maximizing Incremental Information Entropy for Contrastive LearningJiansong Zhang, Zhuoqin Yang, Xu Wu, Xiaoling Luo 等ICLR 2026
- AUC-CL: A Batchsize-Robust Framework for Self-Supervised Contrastive Representation LearningRohan Sharma, Kaiyi Ji, Zhiqiang Xu, Changyou ChenICLR 2024 · 被引用 5 次
- Contrastive Learning with Adversarial ExamplesChih-Hui Ho, Nuno VasconcelosNeurIPS 2020 · 被引用 174 次
- Variational Supervised Contrastive LearningZiwen Wang, Jiajun Fan, Thao Nguyen, Heng Ji 等NeurIPS 2025 · 被引用 7 次
- Modal-aware Bias Constrained Contrastive Learning for Multimodal RecommendationWei Yang, Zhengru Fang, Tianle Zhang, Shiguang Wu 等ACM MM 2023 · 被引用 15 次
