Principal-Agent Bandit Games with Self-Interested and Exploratory Learning Agents
Junyan Liu, Lillian J. Ratliff
摘要
This paper studies the repeated principal-agent bandit game, where the principal indirectly explores an unknown environment by incentivizing an agent to play arms. Unlike prior work that assumes a greedy agent with full knowledge of reward means, we consider a self-interested learning agent who iteratively updates reward estimates and may explore arbitrarily with some probability. As a warm-up, we first consider a self-interested learning agent without exploration. We propose algorithms for both i.i.d. and linear reward settings with bandit feedback in a finite horizon T , achieving regret bounds of O( √ T ) and O(T 2 /3 ), respectively. Specifically, these algorithms rely on a novel elimination framework coupled with new search algorithms which accommodate the uncertainty from the agent's learning behavior. We then extend the framework to handle an exploratory learning agent and develop an algorithm to achieve a O(T 2 /3 ) regret bound in i.i.d. reward setup by enhancing the robustness of our elimination framework to the potential agent exploration. Finally, when our agent model reduces to that in Dogan et al. (2023a), we propose an algorithm based on our robust framework, which achieves a O( √ T ) regret bound, significantly improving upon their O(T 11 /12 ) bound.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Contextual Search in Principal-Agent Games: The Curse of DegeneracyYiding Feng, Mengfan Ma, Bo Peng, Zongqi WanSODA 2026
- Finite-Time Convergence Rates in Stochastic Stackelberg Games with Smooth Algorithmic AgentsEric Frankel, Kshitij Kulkarni, Dmitriy Drusvyatskiy, Sewoong Oh 等ICML 2025
- Learning to Incentivize in Repeated Principal-Agent Problems with Adversarial Agent ArrivalsJunyan Liu, Arnab Maiti, Artin Tajdini, Kevin Jamieson 等ICML 2025
它引用的顶会 Paper5
- Learning with Good Feature Representations in Bandits and in RL with a Generative ModelTor Lattimore, Csaba Szepesvári, Gellért WeiszICML 2020 · 被引用 181 次
- Principal-Agent Reward Shaping in MDPsOmer Ben-Porat, Yishay Mansour, Michal Moshkovitz, Boaz TaitlerAAAI 2024 · 被引用 21 次
- Optimal Contextual Pricing and ExtensionsAllen Liu, Renato Paes Leme, Jon SchneiderSODA 2021 · 被引用 13 次
- Learning to Mitigate Externalities: the Coase Theorem with Hindsight RationalityAntoine Scheid, Aymeric Capitaine, Etienne Boursier, Eric Moulines 等NeurIPS 2024 · 被引用 7 次
- Generalized Principal-Agent Problem with a Learning AgentTao Lin, Yiling ChenICLR 2025
相关 Paper
- Incentivized Learning in Principal-Agent Bandit GamesAntoine Scheid, Daniil Tiapkin, Etienne Boursier, Aymeric Capitaine 等ICML 2024 · 被引用 17 次
- Best Model Identification: A Rested Bandit FormulationLeonardo Cella, Massimiliano Pontil, Claudio GentileICML 2021 · 被引用 6 次
- Contracting with a Learning AgentGuru Guruganesh, Yoav Kolumbus, Jon Schneider, Inbal Talgam-Cohen 等NeurIPS 2024 · 被引用 38 次
- Sequential Information Design: Learning to Persuade in the DarkMartino Bernasconi, Matteo Castiglioni, Alberto Marchesi, Nicola Gatti 等NeurIPS 2022 · 被引用 19 次
- Geometry Meets Incentives: Sample-Efficient Incentivized Exploration with Linear ContextsBen Schiffer, Mark SellkeNeurIPS 2025
