Learning the Optimal Recommendation from Explorative Users
Fan Yao, Chuanhao Li, Denis Nekipelov, Hongning Wang, Haifeng Xu
摘要
We propose a new problem setting to study the sequential interactions between a recommender system and a user. Instead of assuming the user is omniscient, static, and explicit, as the classical practice does, we sketch a more realistic user behavior model, under which the user: 1) rejects recommendations if they are clearly worse than others; 2) updates her utility estimation based on rewards from her accepted recommendations; 3) withholds realized rewards from the system. We formulate the interactions between the system and such an explorative user in a K-armed bandit framework and study the problem of learning the optimal recommendation on the system side. We show that efficient system learning is still possible but is more difficult. In particular, the system can identify the best arm with probability at least 1-delta within O(1/delta) interactions, and we prove this is tight. Our finding contrasts the result for the problem of best arm identification with fixed confidence, in which the best arm can be identified with probability 1-delta within O(log(1/delta)) interactions. This gap illustrates the inevitable cost the system has to pay when it learns from an explorative user's revealed preferences on its recommendations rather than from the realized rewards.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Human vs. Generative AI in Content Creation Competition: Symbiosis or Conflict?Fan Yao, Chuanhao Li, Denis Nekipelov, Hongning Wang 等ICML 2024 · 被引用 31 次
- Unveiling User Satisfaction and Creator Productivity Trade-Offs in Recommendation PlatformsFan Yao, Yiming Liao, Jingzhou Liu, Shaoliang Nie 等NeurIPS 2024 · 被引用 19 次
- Competing for Shareable Arms in Multi-Player Multi-Armed BanditsRenzhe Xu, Haotian Wang, Xingxuan Zhang, Bo Li 等ICML 2023 · 被引用 10 次
- Impact of Decentralized Learning on Player Utilities in Stackelberg GamesKate Donahue, Nicole Immorlica, Meena Jagadeesan, Brendan Lucier 等ICML 2024 · 被引用 9 次
- Learning from a Learning User for Optimal RecommendationsFan Yao, Chuanhao Li, Denis Nekipelov, Hongning Wang 等ICML 2022 · 被引用 8 次
相关 Paper
- Modeling Attrition in Recommender Systems with Departing BanditsOmer Ben-Porat, Lee Cohen, Liu Leqi, Zachary C. Lipton 等AAAI 2022 · 被引用 14 次
- Fiduciary BanditsGal Bahar, Omer Ben-Porat, Kevin Leyton-Brown, Moshe TennenholtzICML 2020 · 被引用 9 次
- Best Arm Identification for Stochastic Rising BanditsMarco Mussi, Alessandro Montenegro, Francesco Trovò, Marcello Restelli 等ICML 2024 · 被引用 4 次
- Covariance-adaptive best arm identificationEl Mehdi Saad, Gilles Blanchard, Nicolas VerzelenNeurIPS 2023 · 被引用 1 次
- Preselection BanditsViktor Bengs, Eyke HüllermeierICML 2020 · 被引用 7 次
