Fiduciary Bandits
Gal Bahar, Omer Ben-Porat, Kevin Leyton-Brown, Moshe Tennenholtz
Abstract
Recommendation systems often face exploration-exploitation tradeoffs: the system can only learn about the desirability of new options by recommending them to some user. Such systems can thus be modeled as multi-armed bandit settings; however, users are self-interested and cannot be made to follow recommendations. We ask whether exploration can nevertheless be performed in a way that scrupulously respects agents' interests---i.e., by a system that acts as a fiduciary. More formally, we introduce a model in which a recommendation system faces an exploration-exploitation tradeoff under the constraint that it can never recommend any action that it knows yields lower reward in expectation than an agent would achieve if it acted alone. Our main contribution is a positive result: an asymptotically optimal, incentive compatible, and ex-ante individually rational recommendation algorithm.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ac4c001d-249b-40d5-8d3a-cf5c99850f2cCited by top-tier papers3
- Modeling Attrition in Recommender Systems with Departing BanditsOmer Ben-Porat, Lee Cohen, Liu Leqi, Zachary C. Lipton et al.AAAI 2022 · 14 citations
- Decongestion by Representation: Learning to Improve Economic Welfare in MarketplacesOmer Nahum, Gali Noti, David C. Parkes, Nir RosenfeldICLR 2024 · 5 citations
- Robust Performance Incentivizing Algorithms for Multi-Armed Bandits with Strategic AgentsSeyed A. Esmaeili, Suho Shin, Aleksandrs SlivkinsAAAI 2025
Related papers
- Incentivizing Combinatorial Bandit ExplorationXinyan Hu, Dung Daniel T. Ngo, Aleksandrs Slivkins, Zhiwei Steven WuNeurIPS 2022 · 14 citations
- Learning the Optimal Recommendation from Explorative UsersFan Yao, Chuanhao Li, Denis Nekipelov, Hongning Wang et al.AAAI 2022 · 8 citations
- Online Certification of Preference-Based Fairness for Personalized Recommender SystemsVirginie Do, Sam Corbett-Davies, Jamal Atif, Nicolas UsunierAAAI 2022 · 47 citations
- Incentivizing Exploration with Linear Contexts and Combinatorial ActionsMark SellkeICML 2023 · 5 citations
- Bandit Learning with Joint Effect of Incentivized Sampling, Delayed Sampling Feedback, and Self-Reinforcing User PreferencesTianchen Zhou, Jia Liu, Chaosheng Dong, Yi SunICLR 2022 · 1 citation
