Control Variates for Slate Off-Policy Evaluation
Nikos Vlassis, Ashok Chandrashekar, Fernando Amat Gil, Nathan Kallus
Abstract
We study the problem of off-policy evaluation from batched contextual bandit data with multidimensional actions, often termed slates. The problem is common to recommender systems and user-interface optimization, and it is particularly challenging because of the combinatorially-sized action space. Swaminathan et al. ( 2017 ) have proposed the pseudoinverse (PI) estimator under the assumption that the conditional mean rewards are additive in actions. Using control variates, we consider a large class of unbiased estimators that includes as specific cases the PI estimator and (asymptotically) its self-normalized variant. By optimizing over this class, we obtain new estimators with risk improvement guarantees over both the PI and the self-normalized PI estimators. Experiments with real-world recommender data as well as synthetic data validate these improvements in practice.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0d50d006-8d8d-4085-8cfa-7c89f533de08Cited by top-tier papers4
- Off-Policy Evaluation for Large Action Spaces via EmbeddingsYuta Saito, Thorsten JoachimsICML 2022 · 62 citations
- Off-Policy Evaluation of Slate Bandit Policies via Optimizing AbstractionHaruka Kiyohara, Masahiro Nomura, Yuta SaitoWWW 2024 · 18 citations
- Exploiting Correlated Auxiliary Feedback in Parameterized BanditsArun Verma, Zhongxiang Dai, Yao Shu, Bryan Kian Hsiang LowNeurIPS 2023 · 6 citations
- Efficient Attention via Control VariatesLin Zheng, Jianbo Yuan, Chong Wang, Lingpeng KongICLR 2023 · 2 citations
Builds on3
- Doubly robust off-policy evaluation with shrinkageYi Su, Maria Dimakopoulou, Akshay Krishnamurthy, Miroslav DudíkICML 2020 · 128 citations
- Counterfactual Evaluation of Slate Recommendations with Sequential Reward InteractionsJames McInerney, Brian Brost, Praveen Chandar, Rishabh Mehrotra et al.KDD 2020 · 45 citations
- Learning from eXtreme Bandit FeedbackRomain Lopez, Inderjit S. Dhillon, Michael I. JordanAAAI 2021 · 26 citations
Related papers
- Distributional Off-Policy Evaluation for Slate RecommendationsShreyas Chaudhari, David Arbour, Georgios Theocharous, Nikos VlassisAAAI 2024 · 2 citations
- Off-Policy Learning in Large Action Spaces: Optimization Matters More Than EstimationImad AOUALI, Otmane SakhiICML 2026
- Empirical Likelihood for Contextual BanditsNikos Karampatziakis, John Langford, Paul MineiroNeurIPS 2020 · 11 citations
- Cross-Domain Off-Policy Evaluation and Learning for Contextual BanditsYuta Natsubori, Masataka Ushiku, Yuta SaitoICLR 2025
- Optimal Off-Policy Evaluation from Multiple Logging PoliciesNathan Kallus, Yuta Saito, Masatoshi UeharaICML 2021 · 44 citations
