Counterfactual Reward Modification for Streaming Recommendation with Delayed Feedback
Xiao Zhang, Haonan Jia, Hanjing Su, Wenhan Wang, Jun Xu, Ji-Rong Wen
Abstract
The user feedbacks could be delayed in many streaming recommendation scenarios. As an example, the user feedbacks to a recommended coupon consist of the immediate feedback on the click event and the delayed feedback on the resultant conversion. The delayed feedbacks pose a challenge of training recommendation models using instances with incomplete labels. When being applied to real products, the challenge becomes more severe as the streaming recommendation models need to be retrained very frequently and the training instances need to be collected over very short time scales. Existing approaches either simply ignore the unobserved feedbacks or heuristically adjust the feedbacks on a static instance set, resulting in biases in the training data and hurting the accuracy of the learned recommenders. In this paper, we propose a novel and theoretic sound counterfactual approach to adjusting the user feedbacks and learning the recommendation models, called CBDF (Counterfactual Bandit with Delayed Feedback). CBDF formulates the streaming recommendation with delayed feedback as a problem of sequential decision making and models it with a batched bandit. To deal with the issue of delayed feedback, at each iteration (episode), a counterfactual importance sampling model is employed to re-weight the original feedbacks and generate the modified rewards. Based on the modified rewards, a batched bandit is learned for conducting online recommendation at the next iteration. Theoretical analysis showed that the modified rewards are statistically unbiased, and the learned bandit policy enjoys a sub-linear regret bound. Experimental results demonstrated that CBDF can outperform the state-of-the-art baselines on a synthetic dataset, the Criteo dataset, and a dataset from Tencent's WeChat app.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get e3b25d11-ce0d-4eb9-8a8b-163eb6241d30Cited by top-tier papers12
- A Model-Agnostic Causal Learning Framework for Recommendation using Search DataZihua Si, Xueran Han, Xiao Zhang, Jun Xu et al.WWW 2022 · 57 citations
- P-MMF: Provider Max-min Fairness Re-ranking in Recommender SystemChen Xu, Sirui Chen, Jun Xu, Weiran Shen et al.WWW 2023 · 41 citations
- PrefRec: Recommender Systems with Human Preferences for Reinforcing Long-term User EngagementWanqi Xue, Qingpeng Cai, Zhenghai Xue, Shuo Sun et al.KDD 2023 · 24 citations
- A Counterfactual Collaborative Session-based Recommender SystemWenzhuo Song, Shoujin Wang, Yan Wang, Kunpeng Liu et al.WWW 2023 · 17 citations
- Reinforcing Long-Term Performance in Recommender Systems with User-Oriented Exploration PolicyChangshuo Zhang, Sirui Chen, Xiao Zhang, Sunhao Dai et al.SIGIR 2024 · 14 citations
Related papers
- Counteracting User Attention Bias in Music Streaming Recommendation via Reward ModificationXiao Zhang, Sunhao Dai, Jun Xu, Zhenhua Dong et al.KDD 2022 · 26 citations
- Capturing Delayed Feedback in Conversion Rate Prediction via Elapsed-Time SamplingJia-Qi Yang, Xiang Li, Shuguang Han, Tao Zhuang et al.AAAI 2021 · 43 citations
- Generalized Delayed Feedback Model with Post-Click Information in Recommender SystemsJia-Qi Yang, De-Chuan ZhanNeurIPS 2022 · 16 citations
- Unbiased Delayed Feedback Label Correction for Conversion Rate PredictionYifan Wang, Peijie Sun, Min Zhang, Qinglin Jia et al.KDD 2023 · 8 citations
- AdaO2B: Adaptive Online to Batch Conversion for Out-of-Distribution GeneralizationXiao Zhang, Sunhao Dai, Jun Xu, Yong Liu et al.AAAI 2025 · 3 citations
