Adaptive Model Design for Markov Decision Process
Siyu Chen, Donglin Yang, Jiayang Li, Senmiao Wang, Zhuoran Yang, Zhaoran Wang
摘要
In a Markov decision process (MDP), an agent interacts with the environment via perceptions and actions. During this process, the agent aims to maximize its own gain. Hence, appropriate regulations are often required, if we hope to take the external costs/benefits of its actions into consideration. In this paper, we study how to regulate such an agent by redesigning model parameters that can affect the rewards and/or the transition kernels. We formulate this problem as a bilevel program, in which the lower-level MDP is regulated by the upper-level model designer. To solve the resulting problem, we develop a scheme that allows the designer to iteratively predict the agent's reaction by solving the MDP and then adaptively update model parameters based on the predicted reaction. The algorithm is first theoretically analyzed and then empirically tested on several MDP models arising in economics and robotics.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Unveiling Induction Heads: Provable Training Dynamics and Feature Learning in TransformersSiyu Chen, Heejune Sheen, Tianhao Wang, Zhuoran YangNeurIPS 2024 · 被引用 48 次
- Contextual Bilevel Reinforcement Learning for Incentive AlignmentVinzenz Thoma, Barna Pásztor, Andreas Krause, Giorgia Ramponi 等NeurIPS 2024 · 被引用 21 次
- Principal-Agent Reward Shaping in MDPsOmer Ben-Porat, Yishay Mansour, Michal Moshkovitz, Boaz TaitlerAAAI 2024 · 被引用 21 次
- Scalable Neural Incentive Design with Parameterized Mean-Field ApproximationNathan Corecco, Batuhan Yardim, Vinzenz Thoma, Zebang Shen 等NeurIPS 2025 · 被引用 1 次
- Learn to change the world: Multi-level reinforcement learning with model-changing actionsZiqing Lu, Babak Hassibi, Lifeng Lai, Weiyu XuICML 2026
它引用的顶会 Paper1
相关 Paper
- Inducing Equilibria via Incentives: Simultaneous Design-and-Play Ensures Global ConvergenceBoyi Liu, Jiayang Li, Zhuoran Yang, Hoi-To Wai 等NeurIPS 2022 · 被引用 29 次
- Learning in Non-Cooperative Configurable Markov Decision ProcessesGiorgia Ramponi, Alberto Maria Metelli, Alessandro Concetti, Marcello RestelliNeurIPS 2021 · 被引用 12 次
- Principled Penalty-based Methods for Bilevel Reinforcement Learning and RLHFHan Shen, Zhuoran Yang, Tianyi ChenICML 2024 · 被引用 35 次
- Planning with Participation ConstraintsHanrui Zhang, Yu Cheng, Vincent ConitzerAAAI 2022 · 被引用 4 次
- Minimally Modifying a Markov Game to Achieve Any Nash Equilibrium and ValueYoung Wu, Jeremy McMahan, Yiding Chen, Yudong Chen 等ICML 2024 · 被引用 3 次
