Regret in Online Recommendation Systems
Kaito Ariu, Narae Ryu, Se-Young Yun, Alexandre Proutière
Abstract
This paper proposes a theoretical analysis of recommendation systems in an online setting, where items are sequentially recommended to users over time. In each round, a user, randomly picked from a population of users, requests a recommendation. The decision-maker observes the user and selects an item from a catalogue of items. Importantly, an item cannot be recommended twice to the same user. The probabilities that a user likes each item are unknown. The performance of the recommendation algorithm is captured through its regret, considering as a reference an Oracle algorithm aware of these probabilities. We investigate various structural assumptions on these probabilities: we derive for each structure regret lower bounds, and devise algorithms achieving these limits. Interestingly, our analysis reveals the relative weights of the different components of regret: the component due to the constraint of not presenting the same item twice to the same user, that due to learning the chances users like items, and finally that arising when learning the underlying structure.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Blocked Collaborative Bandits: Online Collaborative Filtering with Per-Item Budget ConstraintsSoumyabrata Pal, Arun Sai Suggala, Karthikeyan Shanmugam, Prateek JainNeurIPS 2023 · 3 citations
- Multi-User Reinforcement Learning with Low Rank RewardsDheeraj Mysore Nagaraj, Suhas S. Kowshik, Naman Agarwal, Praneeth Netrapalli et al.ICML 2023 · 2 citations
Related papers
- Cascading Bandits: Optimizing Recommendation Frequency in Delayed Feedback EnvironmentsDairui Wang, Junyu Cao, Yan Zhang, Wei QiNeurIPS 2023 · 2 citations
- Online Pricing for Multi-User Multi-Item MarketsYigit Efe Erginbas, Thomas A. Courtade, Kannan Ramchandran, Soham PhadeNeurIPS 2023 · 1 citation
- Online Second Price Auction with Semi-Bandit Feedback under the Non-Stationary SettingHaoyu Zhao, Wei ChenAAAI 2020 · 15 citations
- Fatigue-Aware Bandits for Dependent Click ModelsJunyu Cao, Wei Sun, Zuo-Jun Max Shen, Markus EttlAAAI 2020 · 14 citations
- UniRank: Unimodal Bandit Algorithms for Online RankingCamille-Sovanneary Gauthier, Romaric Gaudel, Élisa FromontICML 2022 · 6 citations
