Bayesian Inferential Risk Evaluation On Multiple IR Systems
Rodger Benham, Ben Carterette, J. Shane Culpepper, Alistair Moffat
Abstract
Information retrieval (IR) ranking models in production systems continually evolve in response to user feedback, insights from research, and new developments. Rather than investing all engineering resources to produce a single challenger to the existing system, a commercial provider might choose to explore multiple new ranking models simultaneously. However, even small changes to a complex model can have unintended consequences. In particular, the per-topic effectiveness profile is likely to change, and even when an overall improvement is achieved, gains are rarely observed for every query, introducing the risk that some users or queries may be negatively impacted by the new model if deployed into production.
Risk adjustments that re-weight losses relative to gains and mitigate such behavior are available when making one-to-one system comparisons, but not for one-to-many or many-to-one comparisons. Moreover, no IR evaluation methodology integrates priors from previous or alternative rankers in a homogeneous inferential framework. In this work, we propose a Bayesian approach where multiple challengers are compared to a single champion. We also show that risk can be incorporated, and demonstrate the benefits of doing so. Finally, the alternative scenario that is commonly encountered in academic research is also considered, when a single challenger is compared against several previous champions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 706e787d-9c91-445d-8303-4c16443d31d5Related papers
- Not All Relevance Scores are Equal: Efficient Uncertainty and Calibration Modeling for Deep Retrieval ModelsDaniel Cohen, Bhaskar Mitra, Oleg Lesota, Navid Rekabsaz et al.SIGIR 2021 · 18 citations
- Offline Evaluation of Ranked Lists using Parametric Estimation of PropensitiesVishwa Vinay, Manoj Kilaru, David ArbourSIGIR 2022
- Topic-oriented Adversarial Attacks against Black-box Neural Ranking ModelsYu-An Liu, Ruqing Zhang, Jiafeng Guo, Maarten de Rijke et al.SIGIR 2023 · 20 citations
- Risk-Sensitive Deep Neural Learning to RankPedro Henrique Silva Rodrigues, Daniel Xavier de Sousa, Thierson Couto Rosa, Marcos André GonçalvesSIGIR 2022 · 7 citations
- BayesCNS: A Unified Bayesian Approach to Address Cold Start and Non-Stationarity in Search Systems at ScaleRandy Ardywibowo, Rakesh Sunki, Shin Tsz Lucy Kuo, Sankalp NayakAAAI 2025
