When Hard Negatives Hurt: Bridging the Generative Discriminative Gap in Hard Negative Synthesis for Retrieval
Zhicheng Zhang, Jiwei Tang, Kuicai Dong, Xiaopeng Li, Jieming Zhu, Jingyu Li, Qianhui Zhu, Fengyuan Lu, Wang Jiaheng, Gang Wang, Hai-Tao Zheng, Zhaocheng Du
Abstract
Hard negative mining has become the dominant strategy for training retrievers, yet it faces intrinsic limitations: negatives are bounded by corpus availability, selected by retriever score rather than diagnostic value, and increasingly contaminated by false positives as the retriever improves. LLM-based synthesis offers a principled alternative, where negatives that are unconstrained, targeted, and free from false positive risk. But we show that naïvely incorporating generated negatives into contrastive learning often degrades retrieval performance. We identify and formalize the root cause as a generative–discriminative gap: LLM generation optimizes for fluent, plausible text, while contrastive learning demands strategic violations of relevance at the decision boundary. Our analysis reveals two compounding failure modes: discriminative-agnostic generation, where the LLM lacks an explicit model of query information needs and defaults to generic or topic-drifted text that provides no contrastive signal; and source-dependent shortcuts, where distributional artifacts enable the model to distinguish negatives by origin rather than relevance, causing gradient drift that actively corrupts optimization. To close this gap, we propose CausalNeg consisting of two main modules: (1) CoT-guided counterfactual perturbation for data construction: decomposes why a document satisfies a query into explicit information requirements, then surgically violates individual requirements to construct negatives with controlled, interpretable hardness. (2) Query-view entropy maximization during training: disperses generated negatives across the similarity spectrum, minimizing the mutual information between source identity and similarity scores to suppress shortcut exploitation. Experiments on 4 retrieval benchmarks show that CausalNeg outperforms mining-only and naïve generation baselines, validating causally grounded synthesis and entropy-regularized training as complementary solutions to the generative–discriminative gap.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on22
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang et al.ICLR 2021 · 1,547 citations
- Contrastive Learning with Hard Negative SamplesJoshua David Robinson, Ching-Yao Chuang, Suvrit Sra, Stefanie JegelkaICLR 2021 · 999 citations
- Optimizing Dense Retrieval Model Training with Hard NegativesJingtao Zhan, Jiaxin Mao, Yiqun Liu, Jiafeng Guo et al.SIGIR 2021 · 242 citations
Related papers
- On Synthetic Data Strategies for Domain-Specific Generative RetrievalHaoyang Wen, Jiang Guo, Yi Zhang, Jiarong Jiang et al.ACL 2025 · 6 citations
- Synthesizing Counterfactual Samples for Effective Image-Text MatchingHao Wei, Shuhui Wang, Xinzhe Han, Zhe Xue et al.ACM MM 2022 · 11 citations
- Constructing Hard-Positive Query-Document Pairs for Dense Retrieval via Phrase RepresentativenessZhanyu Wu, Richong Zhang, Zhijie NieSIGIR 2026
- Relevance Is a Guiding Light: Relevance-aware Adaptive Learning for End-to-end Task-oriented Dialogue SystemZhanpeng Chen, Zhihong Zhu, Wanshi Xu, Xianwei Zhuang et al.EMNLP 2024 · 5 citations
- DocReRank: Single-Page Hard Negative Query Generation for Training Multi-Modal RAG RerankersNavve Wasserman, Oliver Heinimann, Yuval Golbari, Tal Zimbalist et al.EMNLP 2025 · 7 citations
