Adaptively Learning to Select-Rank in Online Platforms
Jingyuan Wang, Perry Dong, Ying Jin, Ruohan Zhan, Zhengyuan Zhou
摘要
Ranking algorithms are fundamental to various online platforms across e-commerce sites to content streaming services. Our research addresses the challenge of adaptively ranking items from a candidate pool for heterogeneous users, a key component in personalizing user experience. We develop a user response model that considers diverse user preferences and the varying effects of item positions, aiming to optimize overall user satisfaction with the ranked list. We frame this problem within a contextual bandits framework, with each ranked list as an action. Our approach incorporates an upper confidence bound to adjust predicted user satisfaction scores and selects the ranking action that maximizes these adjusted scores, efficiently solved via maximum weight imperfect matching. We demonstrate that our algorithm achieves a cumulative regret bound of for ranking out of items in a -dimensional context space over rounds, under the assumption that user responses follow a generalized linear model. This regret alleviates dependence on the ambient action space, whose cardinality grows exponentially with and (thus rendering direct application of existing adaptive learning algorithms -- such as UCB or Thompson sampling -- infeasible). Experiments conducted on both simulated and real-world datasets demonstrate our algorithm outperforms the baseline.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Neural Contextual Bandits with UCB-based ExplorationDongruo Zhou, Lihong Li, Quanquan GuICML 2020 · 被引用 329 次
- Learning a Product Relevance Model from Click-Through Data in E-CommerceShaowei Yao, Jiwei Tan, Xi Chen, Keping Yang 等WWW 2021 · 被引用 48 次
- Mitigating Sentiment Bias for Recommender SystemsChen Lin, Xinyi Liu, Guipeng Xv, Hui LiSIGIR 2021 · 被引用 31 次
- UniRank: Unimodal Bandit Algorithms for Online RankingCamille-Sovanneary Gauthier, Romaric Gaudel, Élisa FromontICML 2022 · 被引用 6 次
相关 Paper
- Optimal Algorithms for Stochastic Contextual Preference BanditsAadirupa SahaNeurIPS 2021 · 被引用 64 次
- Bandits with Ranking FeedbackDavide Maran, Francesco Bacchiocchi, Francesco Emanuele Stradi, Matteo Castiglioni 等NeurIPS 2024 · 被引用 3 次
- Online Clustering of Bandits with Misspecified User ModelsZhiyong Wang, Jize Xie, Xutong Liu, Shuai Li 等NeurIPS 2023 · 被引用 16 次
- Learning from Cross-Modal Behavior Dynamics with Graph-Regularized Neural Contextual BanditXian Wu, Suleyman Cetintas, Deguang Kong, Miao Lu 等WWW 2020 · 被引用 8 次
- Parametric Graph for Unimodal Ranking BanditCamille-Sovanneary Gauthier, Romaric Gaudel, Élisa Fromont, Boammani Aser LompoICML 2021 · 被引用 5 次
