Optimal Robust Subsidy Policies for Irrational Agent in Principal-Agent MDPs
Bowen Hu, Yixin Tao
Abstract
We study a principal-agent problem in a Markov Decision Process where the principal provides subsidies to influence the agent's policy, which in turn determines the accrued rewards. Our focus is on designing a robust subsidy scheme that maximizes the principal’s cumulative expected return, even when the agent displays bounded rationality and may deviate from the optimal action policy after receiving subsidies.
As a baseline, we first analyze the case of a perfectly rational agent and show that the principal’s optimal subsidy coincides with the policy that maximizes social welfare, the sum of the utilities of both the principal and the agent. We then introduce a bounded-rationality model: the globally -incentive-compatible agent, who accepts any policy whose expected cumulative utility lies within of the personal optimum. In this setting, we prove that the optimal robust subsidy scheme problem simplifies to a one-dimensional concave optimization, revealing that optimal subsidies concentrate along social-welfare-maximizing trajectories. We also bound the associated loss in social welfare. Finally, we investigate a finer-grained, state-wise -incentive-compatible model. In this setting, we show that under two natural definitions of state-wise incentive-compatibility, the problem becomes intractable: one definition results in a non-Markovian agent action policy, while the other renders the search for an optimal subsidy scheme NP-hard.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e6852e2f-1a00-453a-a77b-58cb56fa6a16Builds on7
- Bayesian Persuasion in Sequential Decision-MakingJiarui Gan, Rupak Majumdar, Goran Radanovic, Adish SinglaAAAI 2022 · 30 citations
- Contextual Bilevel Reinforcement Learning for Incentive AlignmentVinzenz Thoma, Barna Pásztor, Andreas Krause, Giorgia Ramponi et al.NeurIPS 2024 · 21 citations
- Principal-Agent Reward Shaping in MDPsOmer Ben-Porat, Yishay Mansour, Michal Moshkovitz, Boaz TaitlerAAAI 2024 · 21 citations
- Admissible Policy Teaching through Reward DesignKiarash Banihashem, Adish Singla, Jiarui Gan, Goran RadanovicAAAI 2022 · 18 citations
- Persuading Farsighted Receivers in MDPs: the Power of HonestyMartino Bernasconi, Matteo Castiglioni, Alberto Marchesi, Mirco MuttiNeurIPS 2023 · 8 citations
Related papers
- The Complexity of ContractsPaul Dütting, Tim Roughgarden, Inbal Talgam-CohenSODA 2020 · 26 citations
- Computing Quantal Stackelberg Equilibrium in Extensive-Form GamesJakub Cerný, Viliam Lisý, Branislav Bosanský, Bo AnAAAI 2021 · 7 citations
- Optimal Common Contract with Heterogeneous AgentsShenke Xiao, Zihe Wang, Mengjing Chen, Pingzhong Tang et al.AAAI 2020 · 14 citations
- Learning Optimal Contracts: How to Exploit Small Action SpacesFrancesco Bacchiocchi, Matteo Castiglioni, Alberto Marchesi, Nicola GattiICLR 2024 · 21 citations
- Optimal Mechanism in a Dynamic Stochastic Knapsack EnvironmentJihyeok Jung, Chan-Oi Song, Deok-Joo Lee, Kiho YoonAAAI 2024 · 1 citation
