Offline Evaluation of Ranked Lists using Parametric Estimation of Propensities
Vishwa Vinay, Manoj Kilaru, David Arbour
Abstract
Search engines and recommendation systems attempt to continually improve the quality of the experience they afford to their users. Refining the ranker that produces the lists displayed in response to user requests is an important component of this process. A common practice is for the service providers to make changes (e.g. new ranking features, different ranking models) and A/B test them on a fraction of their users to establish the value of the change. An alternative approach estimates the effectiveness of the proposed changes offline, utilising previously collected clickthrough data on the old ranker to posit what the user behaviour on ranked lists produced by the new ranker would have been. A majority of offline evaluation approaches invoke the well studied inverse propensity weighting to adjust for biases inherent in logged data. In this paper, we propose the use of parametric estimates for these propensities. Specifically, by leveraging well known learning-to-rank methods as subroutines, we show how accurate offline evaluation can be achieved when the new rankings to be evaluated differ from the logged ones.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ef7abbfb-3c49-466b-9c15-1f5ce3195bf7Builds on2
Related papers
- Off-Policy Evaluation for Ranking Policies under Deterministic Logging PoliciesKoichi Tanaka, Kazuki Kawamura, Takanori Muroi, Yusuke Narita et al.ICLR 2026 · 1 citation
- Off-Policy Evaluation of Ranking Policies under Diverse User BehaviorHaruka Kiyohara, Masatoshi Uehara, Yusuke Narita, Nobuyuki Shimizu et al.KDD 2023 · 8 citations
- How do Online Learning to Rank Methods Adapt to Changes of Intent?Shengyao Zhuang, Guido ZucconSIGIR 2021 · 6 citations
- Offline Retrieval Evaluation Without Evaluation MetricsFernando Diaz, Andres FerraroSIGIR 2022 · 8 citations
- Unbiased Learning-to-Rank Needs Unconfounded Propensity EstimationDan Luo, Lixin Zou, Qingyao Ai, Zhiyu Chen et al.SIGIR 2024 · 3 citations
