Not All Relevance Scores are Equal: Efficient Uncertainty and Calibration Modeling for Deep Retrieval Models
Daniel Cohen, Bhaskar Mitra, Oleg Lesota, Navid Rekabsaz, Carsten Eickhoff
Abstract
In any ranking system, the retrieval model outputs a single score for a document based on its belief on how relevant it is to a given search query. While retrieval models have continued to improve with the introduction of increasingly complex architectures, few works have investigated a retrieval model's belief in the score beyond the scope of a single value. We argue that capturing the model's uncertainty with respect to its own scoring of a document is a critical aspect of retrieval that allows for greater use of current models across new document distributions, collections, or even improving effectiveness for down-stream tasks. In this paper, we address this problem via an efficient Bayesian framework for retrieval models which captures the model's belief in the relevance score through a stochastic process while adding only negligible computational overhead. We evaluate this belief via a ranking based calibration metric showing that our approximate Bayesian framework significantly improves a retrieval model's ranking effectiveness through a risk aware reranking as well as its confidence calibration. Lastly, we demonstrate that this additional uncertainty information is actionable and reliable on down-stream tasks represented via cutoff prediction.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cc777791-f969-4646-b6df-2cf6ae4ddaefCited by top-tier papers10
- Joint Multisided Exposure Fairness for RecommendationHaolun Wu, Bhaskar Mitra, Chen Ma, Fernando Diaz et al.SIGIR 2022 · 48 citations
- Can Clicks Be Both Labels and Features?: Unbiased Behavior Feature Collection and Uncertainty-aware Learning to RankTao Yang, Chen Luo, Hanqing Lu, Parth Gupta et al.SIGIR 2022 · 23 citations
- AcuRank: Uncertainty-Aware Adaptive Computation for Listwise RerankingSoyoung Yoon, Gyuwan Kim, Gyu-Hwung Cho, Seung-won HwangNeurIPS 2025 · 15 citations
- Mitigating Exploitation Bias in Learning to Rank with an Uncertainty-aware Empirical Bayes ApproachTao Yang, Cuize Han, Chen Luo, Parth Gupta et al.WWW 2024 · 10 citations
- Stability and Multigroup Fairness in Ranking with Uncertain PredictionsSiddartha Devic, Aleksandra Korolova, David Kempe, Vatsal SharanICML 2024 · 9 citations
Builds on7
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang et al.ICLR 2021 · 1,547 citations
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 1,246 citations
- Bayesian Deep Learning and a Probabilistic Perspective of GeneralizationAndrew Gordon Wilson, Pavel IzmailovNeurIPS 2020 · 845 citations
- Being Bayesian, Even Just a Bit, Fixes Overconfidence in ReLU NetworksAgustinus Kristiadi, Matthias Hein, Philipp HennigICML 2020 · 344 citations
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
Related papers
- REALM: Recursive Relevance Modeling for LLM-based Document Re-RankingPinhuan Wang, Zhiqiu Xia, Chunhua Liao, Feiyi Wang et al.EMNLP 2025
- SGIC: A Self-Guided Iterative Calibration Framework for RAGGuanhua Chen, Yutong Yao, Lidia S. Chao, Xuebo Liu et al.ACL 2025
- On the Calibration and Uncertainty with Pólya-Gamma Augmentation for Dialog Retrieval ModelsTong Ye, Shijing Si, Jianzong Wang, Ning Cheng et al.AAAI 2023 · 3 citations
- Uncertainty Quantification for Retrieval-Augmented ReasoningHeydar Soudani, Hamed Zamani, Faegheh HasibiSIGIR 2026 · 1 citation
- Improving Zero-shot LLM Re-Ranker with Risk MinimizationXiaowei Yuan, Zhao Yang, Yequan Wang, Jun Zhao et al.EMNLP 2024 · 3 citations
