Efficient Training of Retrieval Models using Negative Cache
Erik Lindgren, Sashank J. Reddi, Ruiqi Guo, Sanjiv Kumar
Abstract
Factorized models, such as two tower neural network models, are widely used for scoring (query, document) pairs in information retrieval tasks. These models are typically trained by optimizing the model parameters to score relevant "positive" pairs higher than the irrelevant "negative" ones. While a large set of negatives typically improves the model performance, limited computation and memory budgets place constraints on the number of negatives used during training. In this paper, we develop a novel negative sampling technique for accelerating training with softmax cross-entropy loss. By using cached (possibly stale) item embeddings, our technique enables training with a large pool of negatives with reduced memory and computation. We also develop a streaming variant of our algorithm geared towards very large datasets. Furthermore, we establish a theoretical basis for our approach by showing that updating a very small fraction of the cache at each iteration can still ensure fast convergence. Finally, we experimentally validate our approach and show that it is efficient and compares favorably with more complex, state-of-the-art approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e761dc16-bf97-4ebf-ae12-afdb5829c4e0Cited by top-tier papers10
- Focused Transformer: Contrastive Training for Context ScalingSzymon Tworkowski, Konrad Staniszewski, Mikolaj Pacek, Yuhuai Wu et al.NeurIPS 2023 · 190 citations
- TPU-KNN: K Nearest Neighbor Search at Peak FLOP/sFelix Chern, Blake Hechtman, Andy Davis, Ruiqi Guo et al.NeurIPS 2022 · 31 citations
- Cache-Augmented Inbatch Importance Resampling for Training Recommender RetrieverJin Chen, Defu Lian, Yucheng Li, Baoyun Wang et al.NeurIPS 2022 · 14 citations
- Dual-Encoders for Extreme Multi-label ClassificationNilesh Gupta, Devvrit, Ankit Singh Rawat, Srinadh Bhojanapalli et al.ICLR 2024 · 8 citations
- EMC2: Efficient MCMC Negative Sampling for Contrastive Learning with Global ConvergenceChung-Yiu Yau, Hoi-To Wai, Parameswaran Raman, Soumajyoti Sarkar et al.ICML 2024 · 3 citations
Builds on7
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang et al.ICLR 2021 · 1,547 citations
- Accelerating Large-Scale Inference with Anisotropic Vector QuantizationRuiqi Guo, Philip Sun, Erik Lindgren, Quan Geng et al.ICML 2020 · 539 citations
- Poly-encoders: Architectures and Pre-training Strategies for Fast and Accurate Multi-sentence ScoringSamuel Humeau, Kurt Shuster, Marie-Anne Lachaux, Jason WestonICLR 2020 · 316 citations
- Optimizing Dense Retrieval Model Training with Hard NegativesJingtao Zhan, Jiaxin Mao, Yiqun Liu, Jiafeng Guo et al.SIGIR 2021 · 242 citations
Related papers
- Improving the Accuracy of Dense Retrieval on the Quantized Indexes via Gradient Optimization of the Target EmbeddingsCong Tan, Yongqi Shao, Hong Huo, Tao FangAAAI 2026
- A Fresh Take on Stale Embeddings: Improving Dense Retriever Training with Corrector NetworksNicholas Monath, Will Sussman Grathwohl, Michael Boratko, Rob Fergus et al.ICML 2024 · 1 citation
- A Gradient Accumulation Method for Dense Retriever under Memory ConstraintJaehee Kim, Yukyung Lee, Pilsung KangNeurIPS 2024 · 10 citations
- A Tale of Two Efficient and Informative Negative Sampling DistributionsShabnam Daghaghi, Tharun Medini, Nicholas Meisburger, Beidi Chen et al.ICML 2021 · 11 citations
- Efficiently Teaching an Effective Dense Retriever with Balanced Topic Aware SamplingSebastian Hofstätter, Sheng-Chieh Lin, Jheng-Hong Yang, Jimmy Lin et al.SIGIR 2021 · 297 citations
