Accelerated Convergence for Counterfactual Learning to Rank
Rolf Jagerman, Maarten de Rijke
Abstract
Counterfactual Learning to Rank (LTR) algorithms learn a ranking model from logged user interactions, often collected using a production system. Employing such an offline learning approach has many benefits compared to an online one, but it is challenging as user feedback often contains high levels of bias. Unbiased LTR uses Inverse Propensity Scoring (IPS) to enable unbiased learning from logged user interactions. One of the major difficulties in applying Stochastic Gradient Descent (SGD) approaches to counterfactual learning problems is the large variance introduced by the propensity weights. In this paper we show that the convergence rate of SGD approaches with IPS-weighted gradients suffers from the large variance introduced by the IPS weights: convergence is slow, especially when there are large IPS weights.
To overcome this limitation, we propose a novel learning algorithm, called CounterSample, that has provably better convergence than standard IPS-weighted gradient descent methods. We prove that CounterSample converges faster and complement our theoretical findings with empirical results by performing extensive experimentation in a number of biased LTR scenarios -across optimizers, batch sizes, and different degrees of position bias.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3042306f-9654-41bd-b63c-5f5ce1b4c0dbCited by top-tier papers2
- Adapting Interactional Observation Embedding for Counterfactual Learning to RankMouxiang Chen, Chenghao Liu, Jianling Sun, Steven C. H. HoiSIGIR 2021 · 19 citations
- Learning to Re-rank with Constrained Meta-Optimal TransportAndrés Hoyos IdroboSIGIR 2023 · 1 citation
Related papers
- Policy-Aware Unbiased Learning to Rank for Top-k RankingsHarrie Oosterhuis, Maarten de RijkeSIGIR 2020 · 60 citations
- Safe Deployment for Counterfactual Learning to Rank with Exposure-Based Risk MinimizationShashank Gupta, Harrie Oosterhuis, Maarten de RijkeSIGIR 2023 · 17 citations
- Correcting for Selection Bias in Learning-to-rank SystemsZohreh Ovaisi, Ragib Ahsan, Yifan Zhang, Kathryn Vasilaky et al.WWW 2020 · 123 citations
- Counterfactual Ranking Evaluation with Flexible Click ModelsAlexander Buchholz, Ben London, Giuseppe Di Benedetto, Jan Malte Lichtenberg et al.SIGIR 2024 · 3 citations
- LT2R: Learning to Online Learning to Rank for Web SearchXiaokai Chu, Changying Hao, Shuaiqiang Wang, Dawei Yin et al.ICDE 2024 · 1 citation
