Policy-Aware Unbiased Learning to Rank for Top-k Rankings
Harrie Oosterhuis, Maarten de Rijke
Abstract
Counterfactual Learning to Rank (LTR) methods optimize ranking systems using logged user interactions that contain interaction biases. Existing methods are only unbiased if users are presented with all relevant items in every ranking. There is currently no existing counterfactual unbiased LTR method for top-k rankings. We introduce a novel policy-aware counterfactual estimator for LTR metrics that can account for the effect of a stochastic logging policy. We prove that the policy-aware estimator is unbiased if every relevant item has a non-zero probability to appear in the top-k ranking. Our experimental results show that the performance of our estimator is not affected by the size of k: for any k, the policy-aware estimator reaches the same retrieval performance while learning from top-k feedback as when learning from feedback on the full ranking. Lastly, we introduce novel extensions of traditional LTR methods to perform counterfactual LTR and to optimize top-k metrics. Together, our contributions introduce the first policy-aware unbiased LTR approach that learns from top-k feedback and optimizes top-k metrics. As a result, counterfactual LTR is now applicable to the very prevalent top-k ranking setting in search and recommendation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e1063d73-b9cf-49b2-9fa1-9a173b8b9530Cited by top-tier papers15
- Joint Multisided Exposure Fairness for RecommendationHaolun Wu, Bhaskar Mitra, Chen Ma, Fernando Diaz et al.SIGIR 2022 · 48 citations
- Full Stage Learning to Rank: A Unified Framework for Multi-Stage SystemsKai Zheng, Haijun Zhao, Rui Huang, Beichuan Zhang et al.WWW 2024 · 24 citations
- Can Clicks Be Both Labels and Features?: Unbiased Behavior Feature Collection and Uncertainty-aware Learning to RankTao Yang, Chen Luo, Hanqing Lu, Parth Gupta et al.SIGIR 2022 · 23 citations
- Robust Generalization and Safe Query-Specializationin Counterfactual Learning to RankHarrie Oosterhuis, Maarten de RijkeWWW 2021 · 22 citations
- Adapting Interactional Observation Embedding for Counterfactual Learning to RankMouxiang Chen, Chenghao Liu, Jianling Sun, Steven C. H. HoiSIGIR 2021 · 19 citations
Related papers
- Correcting for Selection Bias in Learning-to-rank SystemsZohreh Ovaisi, Ragib Ahsan, Yifan Zhang, Kathryn Vasilaky et al.WWW 2020 · 123 citations
- Accelerated Convergence for Counterfactual Learning to RankRolf Jagerman, Maarten de RijkeSIGIR 2020 · 13 citations
- Unbiased Learning-to-Rank Needs Unconfounded Propensity EstimationDan Luo, Lixin Zou, Qingyao Ai, Zhiyu Chen et al.SIGIR 2024 · 3 citations
- On the Impact of Outlier Bias on User ClicksFatemeh Sarvi, Ali Vardasbi, Mohammad Aliannejadi, Sebastian Schelter et al.SIGIR 2023 · 6 citations
- LT2R: Learning to Online Learning to Rank for Web SearchXiaokai Chu, Changying Hao, Shuaiqiang Wang, Dawei Yin et al.ICDE 2024 · 1 citation
