Counteracting User Attention Bias in Music Streaming Recommendation via Reward Modification
Xiao Zhang, Sunhao Dai, Jun Xu, Zhenhua Dong, Quanyu Dai, Ji-Rong Wen
摘要
In streaming media applications, like music Apps, songs are recommended in a continuous way in users' daily life. The recommended songs are played automatically although users may not pay any attention to them, posing a challenge of user attention bias in training recommendation models, i.e., the training instances contain a large number of false-positive labels (users' feedback). Existing approaches either directly use the auto-feedbacks or heuristically delete the potential false-positive labels. Both of the approaches lead to biased results because the false-positive labels cause the shift of training data distribution, hurting the accuracy of the recommendation models. In this paper, we propose a learning-based counterfactual approach to adjusting the user auto-feedbacks and learning the recommendation models using Neural Dueling Bandit algorithm, called NDB. Specifically, NDB maintains two neural networks: a user attention network for computing the importance weights that are used for modifying the original rewards, and another random network trained with dueling bandit for conducting online recommendations based on the modified rewards. Theoretical analysis showed that the modified rewards are statistically unbiased, and the learned bandit policy enjoys a sub-linear regret bound. Experimental results demonstrated that NDB can significantly outperform the state-of-the-art baselines.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper7
- P-MMF: Provider Max-min Fairness Re-ranking in Recommender SystemChen Xu, Sirui Chen, Jun Xu, Weiran Shen 等WWW 2023 · 被引用 41 次
- Reinforcing Long-Term Performance in Recommender Systems with User-Oriented Exploration PolicyChangshuo Zhang, Sirui Chen, Xiao Zhang, Sunhao Dai 等SIGIR 2024 · 被引用 14 次
- Reward Imputation with Sketching for Contextual Batched BanditsXiao Zhang, Ninglu Shao, Zihua Si, Jun Xu 等NeurIPS 2023 · 被引用 6 次
- Modeling User Attention in Music RecommendationSunhao Dai, Ninglu Shao, Jieming Zhu, Xiao Zhang 等ICDE 2024 · 被引用 6 次
- AdaO2B: Adaptive Online to Batch Conversion for Out-of-Distribution GeneralizationXiao Zhang, Sunhao Dai, Jun Xu, Yong Liu 等AAAI 2025 · 被引用 3 次
相关 Paper
- Counterfactual Reward Modification for Streaming Recommendation with Delayed FeedbackXiao Zhang, Haonan Jia, Hanjing Su, Wenhan Wang 等SIGIR 2021 · 被引用 60 次
- Online Clustering of Dueling BanditsZhiyong Wang, Jiahang Sun, Mingze Kong, Jize Xie 等ICML 2025
- Neural Dueling Bandits: Preference-Based Optimization with Human FeedbackArun Verma, Zhongxiang Dai, Xiaoqiang Lin, Patrick Jaillet 等ICLR 2025
- Conversational Dueling Bandits in Generalized Linear ModelsShuhua Yang, Hui Yuan, Xiaoying Zhang, Mengdi Wang 等KDD 2024 · 被引用 6 次
- Meta Clustering of Neural BanditsYikun Ban, Yunzhe Qi, Tianxin Wei, Lihui Liu 等KDD 2024 · 被引用 6 次
