Control Variates for Slate Off-Policy Evaluation
Nikos Vlassis, Ashok Chandrashekar, Fernando Amat Gil, Nathan Kallus
摘要
We study the problem of off-policy evaluation from batched contextual bandit data with multidimensional actions, often termed slates. The problem is common to recommender systems and user-interface optimization, and it is particularly challenging because of the combinatorially-sized action space. Swaminathan et al. ( 2017 ) have proposed the pseudoinverse (PI) estimator under the assumption that the conditional mean rewards are additive in actions. Using control variates, we consider a large class of unbiased estimators that includes as specific cases the PI estimator and (asymptotically) its self-normalized variant. By optimizing over this class, we obtain new estimators with risk improvement guarantees over both the PI and the self-normalized PI estimators. Experiments with real-world recommender data as well as synthetic data validate these improvements in practice.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Off-Policy Evaluation for Large Action Spaces via EmbeddingsYuta Saito, Thorsten JoachimsICML 2022 · 被引用 62 次
- Off-Policy Evaluation of Slate Bandit Policies via Optimizing AbstractionHaruka Kiyohara, Masahiro Nomura, Yuta SaitoWWW 2024 · 被引用 18 次
- Exploiting Correlated Auxiliary Feedback in Parameterized BanditsArun Verma, Zhongxiang Dai, Yao Shu, Bryan Kian Hsiang LowNeurIPS 2023 · 被引用 6 次
- Efficient Attention via Control VariatesLin Zheng, Jianbo Yuan, Chong Wang, Lingpeng KongICLR 2023 · 被引用 2 次
它引用的顶会 Paper3
- Doubly robust off-policy evaluation with shrinkageYi Su, Maria Dimakopoulou, Akshay Krishnamurthy, Miroslav DudíkICML 2020 · 被引用 128 次
- Counterfactual Evaluation of Slate Recommendations with Sequential Reward InteractionsJames McInerney, Brian Brost, Praveen Chandar, Rishabh Mehrotra 等KDD 2020 · 被引用 45 次
- Learning from eXtreme Bandit FeedbackRomain Lopez, Inderjit S. Dhillon, Michael I. JordanAAAI 2021 · 被引用 26 次
相关 Paper
- Distributional Off-Policy Evaluation for Slate RecommendationsShreyas Chaudhari, David Arbour, Georgios Theocharous, Nikos VlassisAAAI 2024 · 被引用 2 次
- Off-Policy Learning in Large Action Spaces: Optimization Matters More Than EstimationImad AOUALI, Otmane SakhiICML 2026
- Empirical Likelihood for Contextual BanditsNikos Karampatziakis, John Langford, Paul MineiroNeurIPS 2020 · 被引用 11 次
- Cross-Domain Off-Policy Evaluation and Learning for Contextual BanditsYuta Natsubori, Masataka Ushiku, Yuta SaitoICLR 2025
- Optimal Off-Policy Evaluation from Multiple Logging PoliciesNathan Kallus, Yuta Saito, Masatoshi UeharaICML 2021 · 被引用 44 次
