Optimizing Dense Retrieval Model Training with Hard Negatives
Jingtao Zhan, Jiaxin Mao, Yiqun Liu, Jiafeng Guo, Min Zhang, Shaoping Ma
Abstract
Ranking has always been one of the top concerns in information retrieval researches. For decades, the lexical matching signal has dominated the ad-hoc retrieval process, but solely using this signal in retrieval may cause the vocabulary mismatch problem. In recent years, with the development of representation learning techniques, many researchers turn to Dense Retrieval (DR) models for better ranking performance. Although several existing DR models have already obtained promising results, their performance improvement heavily relies on the sampling of training examples. Many effective sampling strategies are not efficient enough for practical usage, and for most of them, there still lacks theoretical analysis in how and why performance improvement happens. To shed light on these research questions, we theoretically investigate different training strategies for DR models and try to explain why hard negative sampling performs better than random sampling. Through the analysis, we also find that there are many potential risks in static hard negative sampling, which is employed by many existing training methods. Therefore, we propose two training strategies named a Stable Training Algorithm for dense Retrieval (STAR) and a query-side training Algorithm for Directly Optimizing Ranking pErformance (ADORE), respectively. STAR improves the stability of DR training process by introducing random negatives. ADORE replaces the widely-adopted static hard negative sampling method with a dynamic one to directly optimize the ranking performance. Experimental results on two publicly available retrieval benchmark datasets show that either strategy gains significant improvements over existing competitive baselines and a combination of them leads to the best performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fcfc4395-1d6f-4b36-8b40-c962a749e8eeCited by top-tier papers88
- The Power of Noise: Redefining Retrieval for RAG SystemsFlorin Cuconasu, Giovanni Trappolini, Federico Siciliano, Simone Filice et al.SIGIR 2024 · 212 citations
- RocketQAv2: A Joint Training Method for Dense Passage Retrieval and Passage Re-rankingRuiyang Ren, Yingqi Qu, Jing Liu, Wayne Xin Zhao et al.EMNLP 2021 · 147 citations
- Adversarial Retriever-Ranker for Dense Text RetrievalHang Zhang, Yeyun Gong, Yelong Shen, Jiancheng Lv et al.ICLR 2022 · 137 citations
- Multi-View Document Representation Learning for Open-Domain Dense RetrievalShunyu Zhang, Yaobo Liang, Ming Gong, Daxin Jiang et al.ACL 2022 · 80 citations
- Scalable and Effective Generative Information RetrievalHansi Zeng, Chen Luo, Bowen Jin, Sheikh Muhammad Sarwar et al.WWW 2024 · 72 citations
Builds on5
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat et al.ICML 2020 · 2,937 citations
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang et al.ICLR 2021 · 1,547 citations
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 1,246 citations
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu et al.ICLR 2020 · 1,170 citations
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
Related papers
- Constructing Hard-Positive Query-Document Pairs for Dense Retrieval via Phrase RepresentativenessZhanyu Wu, Richong Zhang, Zhijie NieSIGIR 2026
- TriSampler: A Better Negative Sampling Principle for Dense RetrievalZhen Yang, Zhou Shao, Yuxiao Dong, Jie TangAAAI 2024 · 17 citations
- Combining Multiple Supervision for Robust Zero-Shot Dense RetrievalYan Fang, Qingyao Ai, Jingtao Zhan, Yiqun Liu et al.AAAI 2024 · 5 citations
- BERM: Training the Balanced and Extractable Representation for Matching to Improve Generalization Ability of Dense RetrievalShicheng Xu, Liang Pang, Huawei Shen, Xueqi ChengACL 2023 · 8 citations
- CAPSTONE: Curriculum Sampling for Dense Retrieval with Document ExpansionXingwei He, Yeyun Gong, A-Long Jin, Hang Zhang et al.EMNLP 2023 · 3 citations
