Lune

ICML2026Top-tier venue

HOBIT: Hardness Optimized Batch Sampling for InfoNCE Training

Himanshu Dutta, Lokesh Nagalapatti, Yashoteja Prabhu

2026Year

Abstract

Contrastive training with InfoNCE loss and in-batch negatives is the standard approach for learning dual-encoder models. Its effectiveness, however, critically depends on the availability of hard negatives; in their absence, learning quickly saturates. Existing methods address this via explicit hard-negative mining, which is often costly or heuristic-driven. We introduce HOBIT, a principled mini-batch construction method that improves in-batch negative quality by reordering training examples at every epoch. HOBIT\mathrm{\texttt{HOBIT}} solves an optimization problem motivated by the InfoNCE objective to yield mini-batches such that each query in the batch is exposed to hard yet non-contradictory, informative negative examples. We show that the optimization objective is monotone and submodular which in turn leads us to a greedy algorithm that admits the standard O(1−1/e)\mathcal{O}(1 - 1/e) approximation guarantee. Empirically, we show that HOBIT\mathrm{\texttt{HOBIT}} incurs negligible computational overhead while significantly outperforming state-of-the-art batching methods, and remains complementary to existing hard negative mining techniques.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 6eb46800-8295-4ddf-8c1c-2c4e651fa550

Builds on24

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines