Counterfactual Evaluation of Slate Recommendations with Sequential Reward Interactions
James McInerney, Brian Brost, Praveen Chandar, Rishabh Mehrotra, Benjamin A. Carterette
摘要
Users of music streaming, video streaming, news recommendation, and e-commerce services often engage with content in a sequential manner. Providing and evaluating good sequences of recommendations is therefore a central problem for these services. Prior reweighting-based counterfactual evaluation methods either suffer from high variance or make strong independence assumptions about rewards. We propose a new counterfactual estimator that allows for sequential interactions in the rewards with lower variance in an asymptotically unbiased manner. Our method uses graphical assumptions about the causal relationships of the slate to reweight the rewards in the logging policy in a way that approximates the expected sum of rewards under the target policy. Extensive experiments in simulation and on a live recommender system show that our approach outperforms existing methods in terms of bias and data efficiency for the sequential track recommendations problem.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Off-Policy Evaluation for Large Action Spaces via EmbeddingsYuta Saito, Thorsten JoachimsICML 2022 · 被引用 62 次
- Variation Control and Evaluation for Generative Slate RecommendationsShuchang Liu, Fei Sun, Yingqiang Ge, Changhua Pei 等WWW 2021 · 被引用 25 次
- Control Variates for Slate Off-Policy EvaluationNikos Vlassis, Ashok Chandrashekar, Fernando Amat Gil, Nathan KallusNeurIPS 2021 · 被引用 21 次
- Adapting Interactional Observation Embedding for Counterfactual Learning to RankMouxiang Chen, Chenghao Liu, Jianling Sun, Steven C. H. HoiSIGIR 2021 · 被引用 19 次
- Off-Policy Evaluation of Slate Bandit Policies via Optimizing AbstractionHaruka Kiyohara, Masahiro Nomura, Yuta SaitoWWW 2024 · 被引用 18 次
相关 Paper
- Policy-Aware Unbiased Learning to Rank for Top-k RankingsHarrie Oosterhuis, Maarten de RijkeSIGIR 2020 · 被引用 60 次
- Improving Sequential Recommenders through Counterfactual Augmentation of System ExposureZiqi Zhao, Zhaochun Ren, Jiyuan Yang, Zuming Yan 等SIGIR 2025
- Sequential Counterfactual Risk MinimizationHoussam Zenati, Eustache Diemert, Matthieu Martin, Julien Mairal 等ICML 2023 · 被引用 6 次
- Distributional Off-Policy Evaluation for Slate RecommendationsShreyas Chaudhari, David Arbour, Georgios Theocharous, Nikos VlassisAAAI 2024 · 被引用 2 次
- Joint Policy-Value Learning for RecommendationOlivier Jeunen, David Rohde, Flavian Vasile, Martin BompaireKDD 2020 · 被引用 24 次
