Principal-Agent Reward Shaping in MDPs
Omer Ben-Porat, Yishay Mansour, Michal Moshkovitz, Boaz Taitler
摘要
Principal-agent problems arise when one party acts on behalf of another, leading to conflicts of interest. The economic literature has extensively studied principal-agent problems, and recent work has extended this to more complex scenarios such as Markov Decision Processes (MDPs). In this paper, we further explore this line of research by investigating how reward shaping under budget constraints can improve the principal's utility. We study a two-player Stackelberg game where the principal and the agent have different reward functions, and the agent chooses an MDP policy for both players. The principal offers an additional reward to the agent, and the agent picks their policy selfishly to maximize their reward, which is the sum of the original and the offered reward. Our results establish the NP-hardness of the problem and offer polynomial approximation algorithms for two classes of instances: Stochastic trees and deterministic decision processes with a finite horizon.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Contracting with a Learning AgentGuru Guruganesh, Yoav Kolumbus, Jon Schneider, Inbal Talgam-Cohen 等NeurIPS 2024 · 被引用 38 次
- Contextual Bilevel Reinforcement Learning for Incentive AlignmentVinzenz Thoma, Barna Pásztor, Andreas Krause, Giorgia Ramponi 等NeurIPS 2024 · 被引用 21 次
- Incentivized Learning in Principal-Agent Bandit GamesAntoine Scheid, Daniil Tiapkin, Etienne Boursier, Aymeric Capitaine 等ICML 2024 · 被引用 17 次
- Learning to Mitigate Externalities: the Coase Theorem with Hindsight RationalityAntoine Scheid, Aymeric Capitaine, Etienne Boursier, Eric Moulines 等NeurIPS 2024 · 被引用 7 次
- Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game PerspectiveHaichuan Wang, Tao Lin, Lingkai Kong, Ce Li 等ICML 2026 · 被引用 3 次
它引用的顶会 Paper11
- Learning to Utilize Shaping Rewards: A New Approach of Reward ShapingYujing Hu, Weixun Wang, Hangtian Jia, Yixiang Wang 等NeurIPS 2020 · 被引用 256 次
- Scalable Methods for Computing State Similarity in Deterministic Markov Decision ProcessesPablo Samuel CastroAAAI 2020 · 被引用 171 次
- Policy Teaching via Environment Poisoning: Training-time Adversarial Attacks against Reinforcement LearningAmin Rakhsha, Goran Radanovic, Rati Devidze, Xiaojin Zhu 等ICML 2020 · 被引用 145 次
- Consequences of Misaligned AISimon Zhuang, Dylan Hadfield-MenellNeurIPS 2020 · 被引用 120 次
- The Complexity of ContractsPaul Dütting, Tim Roughgarden, Inbal Talgam-CohenSODA 2020 · 被引用 26 次
相关 Paper
- Contract Design Under Approximate Best ResponsesFrancesco Bacchiocchi, Jiarui Gan, Matteo Castiglioni, Alberto Marchesi 等ICML 2025
- Polynomial-Time Approximation Schemes via Utility Alignment: Unit-Demand Pricing and MoreRobin Bowers, Marius Garbea, Emmanouil Pountourakis, Samuel TaggartFOCS 2025 · 被引用 1 次
- Learning in Online Principal-Agent Interactions: The Power of MenusMinbiao Han, Michael Albert, Haifeng XuAAAI 2024 · 被引用 9 次
- Stackelberg Learning with Outcome-based PaymentTom Yan, Chicheng ZhangNeurIPS 2025
- Planning with Participation ConstraintsHanrui Zhang, Yu Cheng, Vincent ConitzerAAAI 2022 · 被引用 4 次
