Robust Generalization and Safe Query-Specializationin Counterfactual Learning to Rank
Harrie Oosterhuis, Maarten de Rijke
摘要
Existing work in counterfactual Learning to Rank (LTR) has focussed on optimizing feature-based models that predict the optimal ranking based on document features. LTR methods based on bandit algorithms often optimize tabular models that memorize the optimal ranking per query. These types of model have their own advantages and disadvantages. Feature-based models provide very robust performance across many queries, including those previously unseen, however, the available features often limit the rankings the model can predict. In contrast, tabular models can converge on any possible ranking through memorization. However, memorization is extremely prone to noise, which makes tabular models reliable only when large numbers of user interactions are available. Can we develop a robust counterfactual LTR method that pursues memorization-based optimization whenever it is safe to do? We introduce the Generalization and Specialization (GENSPEC) algorithm, a robust feature-based counterfactual LTR method that pursues per-query memorization when it is safe to do so. Generalization and Specialization (GENSPEC) optimizes a single feature-based model for generalization: robust performance across all queries, and many tabular models for specialization: each optimized for high performance on a single query. GENSPEC uses novel relative high-confidence bounds to choose which model to deploy per query. By doing so, GENSPEC enjoys the high performance of successfully specialized tabular models with the robustness of a generalized feature-based model. Our results show that GENSPEC leads to optimal performance on queries with sufficient click data, while having robust behavior on queries with little or noisy data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Safe Deployment for Counterfactual Learning to Rank with Exposure-Based Risk MinimizationShashank Gupta, Harrie Oosterhuis, Maarten de RijkeSIGIR 2023 · 被引用 17 次
- Mitigating Exploitation Bias in Learning to Rank with an Uncertainty-aware Empirical Bayes ApproachTao Yang, Cuize Han, Chen Luo, Parth Gupta 等WWW 2024 · 被引用 10 次
- Implicit Feedback for Dense Passage Retrieval: A Counterfactual ApproachShengyao Zhuang, Hang Li, Guido ZucconSIGIR 2022 · 被引用 10 次
- Probabilistic Permutation Graph Search: Black-Box Optimization for Fairness in RankingAli Vardasbi, Fatemeh Sarvi, Maarten de RijkeSIGIR 2022 · 被引用 9 次
它引用的顶会 Paper1
相关 Paper
- Correcting for Selection Bias in Learning-to-rank SystemsZohreh Ovaisi, Ragib Ahsan, Yifan Zhang, Kathryn Vasilaky 等WWW 2020 · 被引用 123 次
- Adapting Interactional Observation Embedding for Counterfactual Learning to RankMouxiang Chen, Chenghao Liu, Jianling Sun, Steven C. H. HoiSIGIR 2021 · 被引用 19 次
- Distributionally Robust Optimization for Unbiased Learning to RankZechun Niu, Lang Mei, Chong Chen, Jiaxin MaoSIGIR 2025
- Unified Off-Policy Learning to Rank: a Reinforcement Learning PerspectiveZeyu Zhang, Yi Su, Hui Yuan, Yiran Wu 等NeurIPS 2023 · 被引用 9 次
- How do Online Learning to Rank Methods Adapt to Changes of Intent?Shengyao Zhuang, Guido ZucconSIGIR 2021 · 被引用 6 次
