Global Selection of Contrastive Batches via Optimization on Sample Permutations
Vin Sachidananda, Ziyi Yang, Chenguang Zhu
摘要
Contrastive Learning has recently achieved stateof-the-art performance in a wide range of unimodal and multimodal tasks. Many contrastive learning approaches use mined hard negatives to make batches more informative during training but these approaches are inefficient as they increase epoch length proportional to the number of mined negatives and require frequent updates of nearest neighbor indices or mining from recent batches. In this work, we provide an alternative to hard negative mining, Global Contrastive Batch Sampling (GCBS), an efficient approximation to the batch assignment problem that upper bounds the gap between the global and training losses, L Global -L T rain , in contrastive learning settings. Through experimentation we find GCBS improves state-of-the-art performance in sentence embedding and code-search tasks. Additionally, GCBS is easy to implement as it requires only a few additional lines of code, does not maintain external data structures such as nearest neighbor indices, is more computationally efficient than the most minimal hard negative mining approaches, and makes no changes to the model being trained. Code is available at https://github.com/vinayak1/GCBS .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Breaking the Batch Barrier (B3) of Contrastive Learning via Smart Batch MiningRaghuveer Thirukovalluru, Rui Meng, Ye Liu, Karthikeyan K 等NeurIPS 2025 · 被引用 30 次
- MoDE: CLIP Data Experts via ClusteringJiawei Ma, Po-Yao Huang, Saining Xie, Shang-Wen Li 等CVPR 2024 · 被引用 8 次
- LogiCoL: Logically-Informed Contrastive Learning for Set-based Dense RetrievalYanzhen Shen, Sihao Chen, Xueqiang Xu, Yunyi Zhang 等EMNLP 2025 · 被引用 1 次
- HOBIT: Hardness Optimized Batch Sampling for InfoNCE TrainingHimanshu Dutta, Lokesh Nagalapatti, Yashoteja PrabhuICML 2026
- Contextual Document EmbeddingsJohn Xavier Morris, Alexander M. RushICLR 2025
它引用的顶会 Paper14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng 等ICLR 2021 · 被引用 1,644 次
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang 等ICLR 2021 · 被引用 1,547 次
相关 Paper
- BatchSampler: Sampling Mini-Batches for Contrastive Learning in Vision, Language, and GraphsZhen Yang, Tinglin Huang, Ming Ding, Yuxiao Dong 等KDD 2023 · 被引用 13 次
- Contrastive Learning with Hard Negative SamplesJoshua David Robinson, Ching-Yao Chuang, Suvrit Sra, Stefanie JegelkaICLR 2021 · 被引用 999 次
- Unsupervised Sentence Representation via Contrastive Learning with Mixing NegativesYanzhao Zhang, Richong Zhang, Samuel Mensah, Xudong Liu 等AAAI 2022 · 被引用 71 次
- B2-Sampling: Fusing Balanced and Biased Sampling for Graph Contrastive LearningMengyue Liu, Yun Lin, Jun Liu, Bohao Liu 等KDD 2023 · 被引用 5 次
- Generating Counterfactual Hard Negative Samples for Graph Contrastive LearningHaoran Yang, Hongxu Chen, Sixiao Zhang, Xiangguo Sun 等WWW 2023 · 被引用 36 次
