Counteracting User Attention Bias in Music Streaming Recommendation via Reward Modification
Xiao Zhang, Sunhao Dai, Jun Xu, Zhenhua Dong, Quanyu Dai, Ji-Rong Wen
Abstract
In streaming media applications, like music Apps, songs are recommended in a continuous way in users' daily life. The recommended songs are played automatically although users may not pay any attention to them, posing a challenge of user attention bias in training recommendation models, i.e., the training instances contain a large number of false-positive labels (users' feedback). Existing approaches either directly use the auto-feedbacks or heuristically delete the potential false-positive labels. Both of the approaches lead to biased results because the false-positive labels cause the shift of training data distribution, hurting the accuracy of the recommendation models. In this paper, we propose a learning-based counterfactual approach to adjusting the user auto-feedbacks and learning the recommendation models using Neural Dueling Bandit algorithm, called NDB. Specifically, NDB maintains two neural networks: a user attention network for computing the importance weights that are used for modifying the original rewards, and another random network trained with dueling bandit for conducting online recommendations based on the modified rewards. Theoretical analysis showed that the modified rewards are statistically unbiased, and the learned bandit policy enjoys a sub-linear regret bound. Experimental results demonstrated that NDB can significantly outperform the state-of-the-art baselines.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get a9e40920-eef3-4823-a681-7d8f51fd8b6eCited by top-tier papers7
- P-MMF: Provider Max-min Fairness Re-ranking in Recommender SystemChen Xu, Sirui Chen, Jun Xu, Weiran Shen et al.WWW 2023 · 41 citations
- Reinforcing Long-Term Performance in Recommender Systems with User-Oriented Exploration PolicyChangshuo Zhang, Sirui Chen, Xiao Zhang, Sunhao Dai et al.SIGIR 2024 · 14 citations
- Reward Imputation with Sketching for Contextual Batched BanditsXiao Zhang, Ninglu Shao, Zihua Si, Jun Xu et al.NeurIPS 2023 · 6 citations
- Modeling User Attention in Music RecommendationSunhao Dai, Ninglu Shao, Jieming Zhu, Xiao Zhang et al.ICDE 2024 · 6 citations
- AdaO2B: Adaptive Online to Batch Conversion for Out-of-Distribution GeneralizationXiao Zhang, Sunhao Dai, Jun Xu, Yong Liu et al.AAAI 2025 · 3 citations
Related papers
- Counterfactual Reward Modification for Streaming Recommendation with Delayed FeedbackXiao Zhang, Haonan Jia, Hanjing Su, Wenhan Wang et al.SIGIR 2021 · 60 citations
- Online Clustering of Dueling BanditsZhiyong Wang, Jiahang Sun, Mingze Kong, Jize Xie et al.ICML 2025
- Neural Dueling Bandits: Preference-Based Optimization with Human FeedbackArun Verma, Zhongxiang Dai, Xiaoqiang Lin, Patrick Jaillet et al.ICLR 2025
- Conversational Dueling Bandits in Generalized Linear ModelsShuhua Yang, Hui Yuan, Xiaoying Zhang, Mengdi Wang et al.KDD 2024 · 6 citations
- Meta Clustering of Neural BanditsYikun Ban, Yunzhe Qi, Tianxin Wei, Lihui Liu et al.KDD 2024 · 6 citations
