HOBIT: Hardness Optimized Batch Sampling for InfoNCE Training
Himanshu Dutta, Lokesh Nagalapatti, Yashoteja Prabhu
Abstract
Contrastive training with InfoNCE loss and in-batch negatives is the standard approach for learning dual-encoder models. Its effectiveness, however, critically depends on the availability of hard negatives; in their absence, learning quickly saturates. Existing methods address this via explicit hard-negative mining, which is often costly or heuristic-driven. We introduce HOBIT, a principled mini-batch construction method that improves in-batch negative quality by reordering training examples at every epoch. solves an optimization problem motivated by the InfoNCE objective to yield mini-batches such that each query in the batch is exposed to hard yet non-contradictory, informative negative examples. We show that the optimization objective is monotone and submodular which in turn leads us to a greedy algorithm that admits the standard approximation guarantee. Empirically, we show that incurs negligible computational overhead while significantly outperforming state-of-the-art batching methods, and remains complementary to existing hard negative mining techniques.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6eb46800-8295-4ddf-8c1c-2c4e651fa550Builds on24
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 2,360 citations
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang et al.ICLR 2021 · 1,547 citations
- Contrastive Learning with Hard Negative SamplesJoshua David Robinson, Ching-Yao Chuang, Suvrit Sra, Stefanie JegelkaICLR 2021 · 999 citations
- Debiased Contrastive LearningChing-Yao Chuang, Joshua Robinson, Yen-Chen Lin, Antonio Torralba et al.NeurIPS 2020 · 761 citations
- Optimizing Dense Retrieval Model Training with Hard NegativesJingtao Zhan, Jiaxin Mao, Yiqun Liu, Jiafeng Guo et al.SIGIR 2021 · 242 citations
Related papers
- Global Selection of Contrastive Batches via Optimization on Sample PermutationsVin Sachidananda, Ziyi Yang, Chenguang ZhuICML 2023 · 6 citations
- When Softmax Fails at the Top: Extreme‑Value Corrections for InfoNCEHasan Sabri Melihcan Erol, Suat Evren, Oktay Ozel, Alexander Morgan et al.ICML 2026
- Empowering Collaborative Filtering with Principled Adversarial Contrastive LossAn Zhang, Leheng Sheng, Zhibo Cai, Xiang Wang et al.NeurIPS 2023 · 56 citations
- Breaking the Batch Barrier (B3) of Contrastive Learning via Smart Batch MiningRaghuveer Thirukovalluru, Rui Meng, Ye Liu, Karthikeyan K et al.NeurIPS 2025 · 30 citations
- EMC2: Efficient MCMC Negative Sampling for Contrastive Learning with Global ConvergenceChung-Yiu Yau, Hoi-To Wai, Parameswaran Raman, Soumajyoti Sarkar et al.ICML 2024 · 3 citations
