Counterfactual Evaluation of Slate Recommendations with Sequential Reward Interactions
James McInerney, Brian Brost, Praveen Chandar, Rishabh Mehrotra, Benjamin A. Carterette
Abstract
Users of music streaming, video streaming, news recommendation, and e-commerce services often engage with content in a sequential manner. Providing and evaluating good sequences of recommendations is therefore a central problem for these services. Prior reweighting-based counterfactual evaluation methods either suffer from high variance or make strong independence assumptions about rewards. We propose a new counterfactual estimator that allows for sequential interactions in the rewards with lower variance in an asymptotically unbiased manner. Our method uses graphical assumptions about the causal relationships of the slate to reweight the rewards in the logging policy in a way that approximates the expected sum of rewards under the target policy. Extensive experiments in simulation and on a live recommender system show that our approach outperforms existing methods in terms of bias and data efficiency for the sequential track recommendations problem.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- Off-Policy Evaluation for Large Action Spaces via EmbeddingsYuta Saito, Thorsten JoachimsICML 2022 · 62 citations
- Variation Control and Evaluation for Generative Slate RecommendationsShuchang Liu, Fei Sun, Yingqiang Ge, Changhua Pei et al.WWW 2021 · 25 citations
- Control Variates for Slate Off-Policy EvaluationNikos Vlassis, Ashok Chandrashekar, Fernando Amat Gil, Nathan KallusNeurIPS 2021 · 21 citations
- Adapting Interactional Observation Embedding for Counterfactual Learning to RankMouxiang Chen, Chenghao Liu, Jianling Sun, Steven C. H. HoiSIGIR 2021 · 19 citations
- Off-Policy Evaluation of Slate Bandit Policies via Optimizing AbstractionHaruka Kiyohara, Masahiro Nomura, Yuta SaitoWWW 2024 · 18 citations
Related papers
- Policy-Aware Unbiased Learning to Rank for Top-k RankingsHarrie Oosterhuis, Maarten de RijkeSIGIR 2020 · 60 citations
- Improving Sequential Recommenders through Counterfactual Augmentation of System ExposureZiqi Zhao, Zhaochun Ren, Jiyuan Yang, Zuming Yan et al.SIGIR 2025
- Sequential Counterfactual Risk MinimizationHoussam Zenati, Eustache Diemert, Matthieu Martin, Julien Mairal et al.ICML 2023 · 6 citations
- Distributional Off-Policy Evaluation for Slate RecommendationsShreyas Chaudhari, David Arbour, Georgios Theocharous, Nikos VlassisAAAI 2024 · 2 citations
- Joint Policy-Value Learning for RecommendationOlivier Jeunen, David Rohde, Flavian Vasile, Martin BompaireKDD 2020 · 24 citations
