Safe Deployment for Counterfactual Learning to Rank with Exposure-Based Risk Minimization
Shashank Gupta, Harrie Oosterhuis, Maarten de Rijke
Abstract
Counterfactual learning to rank (CLTR) relies on exposure-based inverse propensity scoring (IPS), a LTR-specific adaptation of IPS to correct for position bias. While IPS can provide unbiased and consistent estimates, it often suffers from high variance. Especially when little click data is available, this variance can cause CLTR to learn sub-optimal ranking behavior. Consequently, existing CLTR methods bring significant risks with them, as naively deploying their models can result in very negative user experiences.
We introduce a novel risk-aware CLTR method with theoretical guarantees for safe deployment. We apply a novel exposure-based concept of risk regularization to IPS estimation for LTR. Our risk regularization penalizes the mismatch between the ranking behavior of a learned model and a given safe model. Thereby, it ensures that learned ranking models stay close to a trusted model, when there is high uncertainty in IPS estimation, which greatly reduces the risks during deployment. Our experimental results demonstrate the efficacy of our proposed method, which is effective at avoiding initial periods of bad performance when little date is available, while also maintaining high performance at convergence. For the CLTR field, our novel exposure-based risk minimization method enables practitioners to adopt CLTR methods in a safer manner that mitigates many of the risks attached to previous methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9f67e1cd-1a39-4eb4-8b66-391f682816cbCited by top-tier papers1
Ask how each one uses itBuilds on4
- Computationally Efficient Optimization of Plackett-Luce Ranking Models for Relevance and FairnessHarrie OosterhuisSIGIR 2021 · 68 citations
- Policy-Aware Unbiased Learning to Rank for Top-k RankingsHarrie Oosterhuis, Maarten de RijkeSIGIR 2020 · 60 citations
- Policy-Gradient Training of Fair and Unbiased Ranking FunctionsHimank Yadav, Zhengxiao Du, Thorsten JoachimsSIGIR 2021 · 34 citations
- Robust Generalization and Safe Query-Specializationin Counterfactual Learning to RankHarrie Oosterhuis, Maarten de RijkeWWW 2021 · 22 citations
Related papers
- Accelerated Convergence for Counterfactual Learning to RankRolf Jagerman, Maarten de RijkeSIGIR 2020 · 13 citations
- Correcting for Selection Bias in Learning-to-rank SystemsZohreh Ovaisi, Ragib Ahsan, Yifan Zhang, Kathryn Vasilaky et al.WWW 2020 · 123 citations
- Adapting Interactional Observation Embedding for Counterfactual Learning to RankMouxiang Chen, Chenghao Liu, Jianling Sun, Steven C. H. HoiSIGIR 2021 · 19 citations
- On the Impact of Outlier Bias on User ClicksFatemeh Sarvi, Ali Vardasbi, Mohammad Aliannejadi, Sebastian Schelter et al.SIGIR 2023 · 6 citations
- Can Clicks Be Both Labels and Features?: Unbiased Behavior Feature Collection and Uncertainty-aware Learning to RankTao Yang, Chen Luo, Hanqing Lu, Parth Gupta et al.SIGIR 2022 · 23 citations
