Optimal Robust Subsidy Policies for Irrational Agent in Principal-Agent MDPs
Bowen Hu, Yixin Tao
摘要
We study a principal-agent problem in a Markov Decision Process where the principal provides subsidies to influence the agent's policy, which in turn determines the accrued rewards. Our focus is on designing a robust subsidy scheme that maximizes the principal’s cumulative expected return, even when the agent displays bounded rationality and may deviate from the optimal action policy after receiving subsidies.
As a baseline, we first analyze the case of a perfectly rational agent and show that the principal’s optimal subsidy coincides with the policy that maximizes social welfare, the sum of the utilities of both the principal and the agent. We then introduce a bounded-rationality model: the globally -incentive-compatible agent, who accepts any policy whose expected cumulative utility lies within of the personal optimum. In this setting, we prove that the optimal robust subsidy scheme problem simplifies to a one-dimensional concave optimization, revealing that optimal subsidies concentrate along social-welfare-maximizing trajectories. We also bound the associated loss in social welfare. Finally, we investigate a finer-grained, state-wise -incentive-compatible model. In this setting, we show that under two natural definitions of state-wise incentive-compatibility, the problem becomes intractable: one definition results in a non-Markovian agent action policy, while the other renders the search for an optimal subsidy scheme NP-hard.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Bayesian Persuasion in Sequential Decision-MakingJiarui Gan, Rupak Majumdar, Goran Radanovic, Adish SinglaAAAI 2022 · 被引用 30 次
- Contextual Bilevel Reinforcement Learning for Incentive AlignmentVinzenz Thoma, Barna Pásztor, Andreas Krause, Giorgia Ramponi 等NeurIPS 2024 · 被引用 21 次
- Principal-Agent Reward Shaping in MDPsOmer Ben-Porat, Yishay Mansour, Michal Moshkovitz, Boaz TaitlerAAAI 2024 · 被引用 21 次
- Admissible Policy Teaching through Reward DesignKiarash Banihashem, Adish Singla, Jiarui Gan, Goran RadanovicAAAI 2022 · 被引用 18 次
- Persuading Farsighted Receivers in MDPs: the Power of HonestyMartino Bernasconi, Matteo Castiglioni, Alberto Marchesi, Mirco MuttiNeurIPS 2023 · 被引用 8 次
相关 Paper
- The Complexity of ContractsPaul Dütting, Tim Roughgarden, Inbal Talgam-CohenSODA 2020 · 被引用 26 次
- Computing Quantal Stackelberg Equilibrium in Extensive-Form GamesJakub Cerný, Viliam Lisý, Branislav Bosanský, Bo AnAAAI 2021 · 被引用 7 次
- Optimal Common Contract with Heterogeneous AgentsShenke Xiao, Zihe Wang, Mengjing Chen, Pingzhong Tang 等AAAI 2020 · 被引用 14 次
- Learning Optimal Contracts: How to Exploit Small Action SpacesFrancesco Bacchiocchi, Matteo Castiglioni, Alberto Marchesi, Nicola GattiICLR 2024 · 被引用 21 次
- Optimal Mechanism in a Dynamic Stochastic Knapsack EnvironmentJihyeok Jung, Chan-Oi Song, Deok-Joo Lee, Kiho YoonAAAI 2024 · 被引用 1 次
