Learning the Optimal Recommendation from Explorative Users
Fan Yao, Chuanhao Li, Denis Nekipelov, Hongning Wang, Haifeng Xu
Abstract
We propose a new problem setting to study the sequential interactions between a recommender system and a user. Instead of assuming the user is omniscient, static, and explicit, as the classical practice does, we sketch a more realistic user behavior model, under which the user: 1) rejects recommendations if they are clearly worse than others; 2) updates her utility estimation based on rewards from her accepted recommendations; 3) withholds realized rewards from the system. We formulate the interactions between the system and such an explorative user in a K-armed bandit framework and study the problem of learning the optimal recommendation on the system side. We show that efficient system learning is still possible but is more difficult. In particular, the system can identify the best arm with probability at least 1-delta within O(1/delta) interactions, and we prove this is tight. Our finding contrasts the result for the problem of best arm identification with fixed confidence, in which the best arm can be identified with probability 1-delta within O(log(1/delta)) interactions. This gap illustrates the inevitable cost the system has to pay when it learns from an explorative user's revealed preferences on its recommendations rather than from the realized rewards.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 94db45c2-c516-447f-9d03-e5dcd578f455Cited by top-tier papers8
- Human vs. Generative AI in Content Creation Competition: Symbiosis or Conflict?Fan Yao, Chuanhao Li, Denis Nekipelov, Hongning Wang et al.ICML 2024 · 31 citations
- Unveiling User Satisfaction and Creator Productivity Trade-Offs in Recommendation PlatformsFan Yao, Yiming Liao, Jingzhou Liu, Shaoliang Nie et al.NeurIPS 2024 · 19 citations
- Competing for Shareable Arms in Multi-Player Multi-Armed BanditsRenzhe Xu, Haotian Wang, Xingxuan Zhang, Bo Li et al.ICML 2023 · 10 citations
- Impact of Decentralized Learning on Player Utilities in Stackelberg GamesKate Donahue, Nicole Immorlica, Meena Jagadeesan, Brendan Lucier et al.ICML 2024 · 9 citations
- Learning from a Learning User for Optimal RecommendationsFan Yao, Chuanhao Li, Denis Nekipelov, Hongning Wang et al.ICML 2022 · 8 citations
Related papers
- Modeling Attrition in Recommender Systems with Departing BanditsOmer Ben-Porat, Lee Cohen, Liu Leqi, Zachary C. Lipton et al.AAAI 2022 · 14 citations
- Fiduciary BanditsGal Bahar, Omer Ben-Porat, Kevin Leyton-Brown, Moshe TennenholtzICML 2020 · 9 citations
- Best Arm Identification for Stochastic Rising BanditsMarco Mussi, Alessandro Montenegro, Francesco Trovò, Marcello Restelli et al.ICML 2024 · 4 citations
- Covariance-adaptive best arm identificationEl Mehdi Saad, Gilles Blanchard, Nicolas VerzelenNeurIPS 2023 · 1 citation
- Preselection BanditsViktor Bengs, Eyke HüllermeierICML 2020 · 7 citations
