Effective Contrastive Weighting for Dense Query Expansion
Xiao Wang, Sean MacAvaney, Craig Macdonald, Iadh Ounis
Abstract
Verbatim queries submitted to search engines often do not sufficiently describe the user’s search intent. Pseudo-relevance feedback (PRF) techniques, which modify a query’s representation using the top-ranked documents, have been shown to overcome such inadequacies and improve retrieval effectiveness for both lexical methods (e.g., BM25) and dense methods (e.g., ANCE, ColBERT). For instance, the recent ColBERT-PRF approach heuristically chooses new embeddings to add to the query representation using the inverse document frequency (IDF) of the underlying tokens. However, this heuristic potentially ignores the valuable context encoded by the embeddings. In this work, we present a contrastive solution that learns to select the most useful embeddings for expansion. More specifically, a deep language model-based contrastive weighting model, called CWPRF, is trained to learn to discriminate between relevant and non-relevant documents for semantic search. Our experimental results show that our contrastive weighting model can aid to select useful expansion embeddings and outperform various baselines. In particular, CWPRF can improve nDCG@10 by up to to 4.1% compared to an existing PRF approach for ColBERT while maintaining its efficiency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 062272ca-c0a2-417e-84cb-e78013f9de5cCited by top-tier papers2
- ConceptCarve: Dynamic Realization of EvidenceEylon Caplan, Dan GoldwasserACL 2025
- PLAID-PRF: Pseudo-Relevance Feedback with Centroid-like Tokens in PLAIDXiao Wang, Sean MacAvaney, Craig MacdonaldSIGIR 2026
Builds on6
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang et al.ICLR 2021 · 1,547 citations
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 1,246 citations
- Optimizing Dense Retrieval Model Training with Hard NegativesJingtao Zhan, Jiaxin Mao, Yiqun Liu, Jiafeng Guo et al.SIGIR 2021 · 242 citations
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
- Context-Aware Document Term Weighting for Ad-Hoc SearchZhuyun Dai, Jamie CallanWWW 2020 · 123 citations
Related papers
- Leveraging Cognitive Complexity of Texts for Contextualization in Dense RetrievalEffrosyni Sokli, Georgios Peikos, Pranav Kasela, Gabriella PasiEMNLP 2025
- Constructing Hard-Positive Query-Document Pairs for Dense Retrieval via Phrase RepresentativenessZhanyu Wu, Richong Zhang, Zhijie NieSIGIR 2026
- LoL: A Comparative Regularization Loss over Query Reformulation Losses for Pseudo-Relevance FeedbackYunchang Zhu, Liang Pang, Yanyan Lan, Huawei Shen et al.SIGIR 2022 · 6 citations
- DOGR: Leveraging Document-Oriented Contrastive Learning in Generative RetrievalPenghao Lu, Xin Dong, Yuansheng Zhou, Lei Cheng et al.AAAI 2025
- Decoding a Neural Retriever's Latent Space for Query SuggestionLeonard Adolphs, Michelle Chen Huebscher, Christian Buck, Sertan Girgin et al.EMNLP 2022 · 4 citations
