Fiduciary Bandits
Gal Bahar, Omer Ben-Porat, Kevin Leyton-Brown, Moshe Tennenholtz
摘要
Recommendation systems often face exploration-exploitation tradeoffs: the system can only learn about the desirability of new options by recommending them to some user. Such systems can thus be modeled as multi-armed bandit settings; however, users are self-interested and cannot be made to follow recommendations. We ask whether exploration can nevertheless be performed in a way that scrupulously respects agents' interests---i.e., by a system that acts as a fiduciary. More formally, we introduce a model in which a recommendation system faces an exploration-exploitation tradeoff under the constraint that it can never recommend any action that it knows yields lower reward in expectation than an agent would achieve if it acted alone. Our main contribution is a positive result: an asymptotically optimal, incentive compatible, and ex-ante individually rational recommendation algorithm.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Modeling Attrition in Recommender Systems with Departing BanditsOmer Ben-Porat, Lee Cohen, Liu Leqi, Zachary C. Lipton 等AAAI 2022 · 被引用 14 次
- Decongestion by Representation: Learning to Improve Economic Welfare in MarketplacesOmer Nahum, Gali Noti, David C. Parkes, Nir RosenfeldICLR 2024 · 被引用 5 次
- Robust Performance Incentivizing Algorithms for Multi-Armed Bandits with Strategic AgentsSeyed A. Esmaeili, Suho Shin, Aleksandrs SlivkinsAAAI 2025
相关 Paper
- Incentivizing Combinatorial Bandit ExplorationXinyan Hu, Dung Daniel T. Ngo, Aleksandrs Slivkins, Zhiwei Steven WuNeurIPS 2022 · 被引用 14 次
- Learning the Optimal Recommendation from Explorative UsersFan Yao, Chuanhao Li, Denis Nekipelov, Hongning Wang 等AAAI 2022 · 被引用 8 次
- Online Certification of Preference-Based Fairness for Personalized Recommender SystemsVirginie Do, Sam Corbett-Davies, Jamal Atif, Nicolas UsunierAAAI 2022 · 被引用 47 次
- Incentivizing Exploration with Linear Contexts and Combinatorial ActionsMark SellkeICML 2023 · 被引用 5 次
- Bandit Learning with Joint Effect of Incentivized Sampling, Delayed Sampling Feedback, and Self-Reinforcing User PreferencesTianchen Zhou, Jia Liu, Chaosheng Dong, Yi SunICLR 2022 · 被引用 1 次
