BatchSampler: Sampling Mini-Batches for Contrastive Learning in Vision, Language, and Graphs
Zhen Yang, Tinglin Huang, Ming Ding, Yuxiao Dong, Rex Ying, Yukuo Cen, Yangliao Geng, Jie Tang
摘要
In-Batch contrastive learning is a state-of-the-art self-supervised method that brings semantically-similar instances close while pushing dissimilar instances apart within a mini-batch. Its key to success is the negative sharing strategy, in which every instance serves as a negative for the others within the mini-batch. Recent studies aim to improve performance by sampling hard negatives within the current mini-batch, whose quality is bounded by the mini-batch itself. In this work, we propose to improve contrastive learning by sampling mini-batches from the input data. We present BatchSamplercode is available at BatchSampler to sample mini-batches of hard-to-distinguish (i.e., hard and true negatives to each other) instances. To make each mini-batch have fewer false negatives, we design the proximity graph of randomly-selected instances. To form the mini-batch, we leverage random walk with restart on the proximity graph to help sample hard-to-distinguish instances. BatchSampler is a simple and general technique that can be directly plugged into existing contrastive learning models in vision, language, and graphs. Extensive experiments on datasets of three modalities show that BatchSampler can consistently improve the performance of powerful contrastive models, as shown by significant improvements of SimCLR on ImageNet-100, SimCSE on STS (language), and GraphCL and MVGRL on graph datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Breaking the Batch Barrier (B3) of Contrastive Learning via Smart Batch MiningRaghuveer Thirukovalluru, Rui Meng, Ye Liu, Karthikeyan K 等NeurIPS 2025 · 被引用 30 次
- LLMEmb: Large Language Model Can Be a Good Embedding Generator for Sequential RecommendationQidong Liu, Xian Wu, Wanyu Wang, Yejing Wang 等AAAI 2025 · 被引用 12 次
- Bypassing Skip-Gram Negative Sampling: Dimension Regularization as a More Efficient Alternative for Graph EmbeddingsDavid Liu, Arjun Seshadri, Tina Eliassi-Rad, Johan UganderKDD 2025
- HOBIT: Hardness Optimized Batch Sampling for InfoNCE TrainingHimanshu Dutta, Lokesh Nagalapatti, Yashoteja PrabhuICML 2026
它引用的顶会 Paper28
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal 等NeurIPS 2020 · 被引用 5,249 次
- Graph Contrastive Learning with AugmentationsYuning You, Tianlong Chen, Yongduo Sui, Ting Chen 等NeurIPS 2020 · 被引用 3,042 次
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
- An Empirical Study of Training Self-Supervised Vision TransformersXinlei Chen, Saining Xie, Kaiming HeICCV 2021 · 被引用 2,340 次
相关 Paper
- Global Selection of Contrastive Batches via Optimization on Sample PermutationsVin Sachidananda, Ziyi Yang, Chenguang ZhuICML 2023 · 被引用 6 次
- Generating Counterfactual Hard Negative Samples for Graph Contrastive LearningHaoran Yang, Hongxu Chen, Sixiao Zhang, Xiangguo Sun 等WWW 2023 · 被引用 36 次
- Contrastive Learning with Hard Negative SamplesJoshua David Robinson, Ching-Yao Chuang, Suvrit Sra, Stefanie JegelkaICLR 2021 · 被引用 999 次
- B2-Sampling: Fusing Balanced and Biased Sampling for Graph Contrastive LearningMengyue Liu, Yun Lin, Jun Liu, Bohao Liu 等KDD 2023 · 被引用 5 次
- Unsupervised Sentence Representation via Contrastive Learning with Mixing NegativesYanzhao Zhang, Richong Zhang, Samuel Mensah, Xudong Liu 等AAAI 2022 · 被引用 71 次
