Incentivized Learning in Principal-Agent Bandit Games
Antoine Scheid, Daniil Tiapkin, Etienne Boursier, Aymeric Capitaine, Eric Moulines, Michael I. Jordan, El-Mahdi El-Mhamdi, Alain Oliviero Durmus
Abstract
This work considers a repeated principal-agent bandit game, where the principal can only interact with her environment through the agent. The principal and the agent have misaligned objectives and the choice of action is only left to the agent. However, the principal can influence the agent's decisions by offering incentives which add up to his rewards. The principal aims to iteratively learn an incentive policy to maximize her own total utility. This framework extends usual bandit problems and is motivated by several practical applications, such as healthcare or ecological taxation, where traditionally used mechanism design theories often overlook the learning aspect of the problem. We present nearly optimal (with respect to a horizon ) learning algorithms for the principal's regret in both multi-armed and linear contextual settings. Finally, we support our theoretical guarantees through numerical experiments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 629375aa-3f77-4ca5-90f3-c655b927a404Cited by top-tier papers6
- Learning to Mitigate Externalities: the Coase Theorem with Hindsight RationalityAntoine Scheid, Aymeric Capitaine, Etienne Boursier, Eric Moulines et al.NeurIPS 2024 · 7 citations
- The Optimal Sample Complexity of Linear ContractsMikael Moller HogsgaardICML 2026 · 2 citations
- Stackelberg Learning with Outcome-based PaymentTom Yan, Chicheng ZhangNeurIPS 2025
- Provably Efficient Algorithm for Best Scoring Rule Identification in Online Principal-Agent Information AcquisitionZichen Wang, Chuanhao Li, Huazheng WangICML 2025
- Contextual Search in Principal-Agent Games: The Curse of DegeneracyYiding Feng, Mengfan Ma, Bo Peng, Zongqi WanSODA 2026
Builds on4
- Nearly Optimal Algorithms for Linear Contextual Bandits with Adversarial CorruptionsJiafan He, Dongruo Zhou, Tong Zhang, Quanquan GuNeurIPS 2022 · 66 citations
- Online Bayesian PersuasionMatteo Castiglioni, Andrea Celli, Alberto Marchesi, Nicola GattiNeurIPS 2020 · 26 citations
- Principal-Agent Reward Shaping in MDPsOmer Ben-Porat, Yishay Mansour, Michal Moshkovitz, Boaz TaitlerAAAI 2024 · 21 citations
- Incentivizing Combinatorial Bandit ExplorationXinyan Hu, Dung Daniel T. Ngo, Aleksandrs Slivkins, Zhiwei Steven WuNeurIPS 2022 · 14 citations
Related papers
- Principal-Agent Bandit Games with Self-Interested and Exploratory Learning AgentsJunyan Liu, Lillian J. RatliffICML 2025
- Learning to Incentivize in Repeated Principal-Agent Problems with Adversarial Agent ArrivalsJunyan Liu, Arnab Maiti, Artin Tajdini, Kevin Jamieson et al.ICML 2025
- Contracting with a Learning AgentGuru Guruganesh, Yoav Kolumbus, Jon Schneider, Inbal Talgam-Cohen et al.NeurIPS 2024 · 38 citations
- Generalized Principal-Agent Problem with a Learning AgentTao Lin, Yiling ChenICLR 2025
- Learning a Game by Paying the AgentsBrian Hu Zhang, Tao Lin, Yiling Chen, Tuomas SandholmICLR 2026 · 1 citation
