Modeling Attrition in Recommender Systems with Departing Bandits
Omer Ben-Porat, Lee Cohen, Liu Leqi, Zachary C. Lipton, Yishay Mansour
摘要
Traditionally, when recommender systems are formalized as multi-armed bandits, the policy of the recommender system influences the rewards accrued, but not the length of interaction. However, in real-world systems, dissatisfied users may depart (and never come back). In this work, we propose a novel multi-armed bandit setup that captures such policydependent horizons. Our setup consists of a finite set of user types, and multiple arms with Bernoulli payoffs. Each (user type, arm) tuple corresponds to an (unknown) reward probability. Each user's type is initially unknown and can only be inferred through their response to recommendations. Moreover, if a user is dissatisfied with their recommendation, they might depart the system. We first address the case where all users share the same type, demonstrating that a recent UCBbased algorithm is optimal. We then move forward to the more challenging case, where users are divided among two types. While naive approaches cannot handle this setting, we provide an efficient learning algorithm that achieves Õ( √ T ) regret, where T is the number of users. 1 We denote by [n] the set 1, . . . , n. 2 We formalize the reward as is standard in the online learning literature, from the perspective of the learner. However, defining the reward from the user perspective by, e.g., considering her utility as the number of clicks she gives or the number of articles she reads induces the same model.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Learning with Exposure Constraints in Recommendation SystemsOmer Ben-Porat, Rotem TorkanWWW 2023 · 被引用 16 次
- Classification Under Strategic Self-SelectionGuy Horowitz, Yonatan Sommer, Moran Koren, Nir RosenfeldICML 2024 · 被引用 8 次
- Direct Routing Gradient (DRGrad): A Personalized Information Surgery for Multi-Task Learning (MTL) RecommendationsYuguang Liu, Yiyun Miao, Luyao XiaAAAI 2025 · 被引用 2 次
它引用的顶会 Paper4
- Achieving Fairness in the Stochastic Multi-Armed Bandit ProblemVishakha Patil, Ganesh Ghalme, Vineet Nair, Y. NarahariAAAI 2020 · 被引用 131 次
- Rebounding Bandits for Modeling Satiation EffectsLiu Leqi, Fatma Kilinç-Karzan, Zachary C. Lipton, Alan L. MontgomeryNeurIPS 2021 · 被引用 30 次
- Fatigue-Aware Bandits for Dependent Click ModelsJunyu Cao, Wei Sun, Zuo-Jun Max Shen, Markus EttlAAAI 2020 · 被引用 14 次
- Fiduciary BanditsGal Bahar, Omer Ben-Porat, Kevin Leyton-Brown, Moshe TennenholtzICML 2020 · 被引用 9 次
相关 Paper
- Learning the Optimal Recommendation from Explorative UsersFan Yao, Chuanhao Li, Denis Nekipelov, Hongning Wang 等AAAI 2022 · 被引用 8 次
- Bandit Learning with Joint Effect of Incentivized Sampling, Delayed Sampling Feedback, and Self-Reinforcing User PreferencesTianchen Zhou, Jia Liu, Chaosheng Dong, Yi SunICLR 2022 · 被引用 1 次
- Incentivized Bandit Learning with Self-Reinforcing User PreferencesTianchen Zhou, Jia Liu, Chaosheng Dong, Jingyuan DengICML 2021 · 被引用 2 次
- Cascading Bandits: Optimizing Recommendation Frequency in Delayed Feedback EnvironmentsDairui Wang, Junyu Cao, Yan Zhang, Wei QiNeurIPS 2023 · 被引用 2 次
- Dynamic Planning and Learning under Recovering RewardsDavid Simchi-Levi, Zeyu Zheng, Feng ZhuICML 2021 · 被引用 6 次
