Bandit Social Leaning Dynamics with Exploration Episodes
Kiarash Banihashem, Natalie Collina, Alex Slivkins
Abstract
We study a stylized social learning dynamics where self-interested agents collectively follow a simple multi-armed bandit protocol. Each agent controls an "episode": a short sequence of consecutive decisions. Motivating applications include users repeatedly interacting with an AI, or repeatedly shopping at a marketplace. While agents are incentivized to explore within their respective episodes, we show that the aggregate exploration fails: e.g., its Bayesian regret grows linearly over time. In fact, such failure is a (very) typical case, not just a worst-case scenario. This conclusion persists even if an agent's per-episode utility is some fixed function of the per-round outcomes: e.g., or , not just the sum. Thus, externally driven exploration is needed even when some amount of exploration happens organically.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on2
- Unreasonable Effectiveness of Greedy Algorithms in Multi-Armed Bandit with Many ArmsMohsen Bayati, Nima Hamidi, Ramesh Johari, Khashayar KhosraviNeurIPS 2020 · 10 citations
- Bandit Social Learning under Myopic BehaviorKiarash Banihashem, MohammadTaghi Hajiaghayi, Suho Shin, Aleksandrs SlivkinsNeurIPS 2023 · 6 citations
Related papers
- (Almost) Free Incentivized Exploration from Decentralized Learning AgentsChengshuai Shi, Haifeng Xu, Wei Xiong, Cong ShenNeurIPS 2021 · 10 citations
- Impact of Decentralized Learning on Player Utilities in Stackelberg GamesKate Donahue, Nicole Immorlica, Meena Jagadeesan, Brendan Lucier et al.ICML 2024 · 9 citations
- Fiduciary BanditsGal Bahar, Omer Ben-Porat, Kevin Leyton-Brown, Moshe TennenholtzICML 2020 · 9 citations
- The Impact of Exploration on Convergence and Performance of Multi-Agent Q-Learning DynamicsAamal Abbas Hussain, Francesco Belardinelli, Dario PaccagnanICML 2023 · 2 citations
- Incentivized Learning in Principal-Agent Bandit GamesAntoine Scheid, Daniil Tiapkin, Etienne Boursier, Aymeric Capitaine et al.ICML 2024 · 17 citations
