Lune

NeurIPS2025顶会

Stochastic Principal-Agent Problems: Computing and Learning Optimal History-Dependent Policies

Jiarui Gan, Rupak Majumdar, Debmalya Mandal, Goran Radanovic

2025年份

摘要

We study a stochastic principal-agent model. A principal and an agent interact in a stochastic environment, each privy to observations about the state not available to the other. The principal has the power of commitment, both to elicit information from the agent and to signal her own information. The players communicate with each other and then select actions independently. Both players are far-sighted , aiming to maximize their total payoffs over the entire time horizon. We consider both the computation and learning of the principal’s optimal policy. The key challenge lies in enabling history-dependent policies, which are essential for achieving optimality in this model but difficult to cope with because of the exponential growth of possible histories as the size of the model increases; explicit representation of history-dependent policies is infeasible as a result. To address this challenge, we develop algorithmic techniques based on the concept of inducible value set . The techniques yield an efficient algorithm that computes an ϵ -approximate optimal policy in time polynomial in 1 /ϵ . We also present an efficient learning algorithm for an episodic reinforcement learning setting with unknown transition probabilities. The algorithm achieves sublinear regret (cid:101) O ( T 2 / 3 ) for both players over T episodes.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper12

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖