Robust Generalization and Safe Query-Specializationin Counterfactual Learning to Rank
Harrie Oosterhuis, Maarten de Rijke
Abstract
Existing work in counterfactual Learning to Rank (LTR) has focussed on optimizing feature-based models that predict the optimal ranking based on document features. LTR methods based on bandit algorithms often optimize tabular models that memorize the optimal ranking per query. These types of model have their own advantages and disadvantages. Feature-based models provide very robust performance across many queries, including those previously unseen, however, the available features often limit the rankings the model can predict. In contrast, tabular models can converge on any possible ranking through memorization. However, memorization is extremely prone to noise, which makes tabular models reliable only when large numbers of user interactions are available. Can we develop a robust counterfactual LTR method that pursues memorization-based optimization whenever it is safe to do? We introduce the Generalization and Specialization (GENSPEC) algorithm, a robust feature-based counterfactual LTR method that pursues per-query memorization when it is safe to do so. Generalization and Specialization (GENSPEC) optimizes a single feature-based model for generalization: robust performance across all queries, and many tabular models for specialization: each optimized for high performance on a single query. GENSPEC uses novel relative high-confidence bounds to choose which model to deploy per query. By doing so, GENSPEC enjoys the high performance of successfully specialized tabular models with the robustness of a generalized feature-based model. Our results show that GENSPEC leads to optimal performance on queries with sufficient click data, while having robust behavior on queries with little or noisy data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d656e381-20da-4cdc-95d0-998438cce059Cited by top-tier papers4
- Safe Deployment for Counterfactual Learning to Rank with Exposure-Based Risk MinimizationShashank Gupta, Harrie Oosterhuis, Maarten de RijkeSIGIR 2023 · 17 citations
- Mitigating Exploitation Bias in Learning to Rank with an Uncertainty-aware Empirical Bayes ApproachTao Yang, Cuize Han, Chen Luo, Parth Gupta et al.WWW 2024 · 10 citations
- Implicit Feedback for Dense Passage Retrieval: A Counterfactual ApproachShengyao Zhuang, Hang Li, Guido ZucconSIGIR 2022 · 10 citations
- Probabilistic Permutation Graph Search: Black-Box Optimization for Fairness in RankingAli Vardasbi, Fatemeh Sarvi, Maarten de RijkeSIGIR 2022 · 9 citations
Builds on1
Related papers
- Correcting for Selection Bias in Learning-to-rank SystemsZohreh Ovaisi, Ragib Ahsan, Yifan Zhang, Kathryn Vasilaky et al.WWW 2020 · 123 citations
- Adapting Interactional Observation Embedding for Counterfactual Learning to RankMouxiang Chen, Chenghao Liu, Jianling Sun, Steven C. H. HoiSIGIR 2021 · 19 citations
- Distributionally Robust Optimization for Unbiased Learning to RankZechun Niu, Lang Mei, Chong Chen, Jiaxin MaoSIGIR 2025
- Unified Off-Policy Learning to Rank: a Reinforcement Learning PerspectiveZeyu Zhang, Yi Su, Hui Yuan, Yiran Wu et al.NeurIPS 2023 · 9 citations
- How do Online Learning to Rank Methods Adapt to Changes of Intent?Shengyao Zhuang, Guido ZucconSIGIR 2021 · 6 citations
