Lune

NeurIPS2025顶会

Scalable Neural Incentive Design with Parameterized Mean-Field Approximation

Nathan Corecco, Batuhan Yardim, Vinzenz Thoma, Zebang Shen, Niao He

2025年份
1被引次数
1顶会引用

摘要

Designing incentives for a multi-agent system to induce a desirable Nash equilibrium is both a crucial and challenging problem appearing in many decision-making domains, especially for a large number of agents N . Under the exchangeability assumption, we formalize this incentive design (ID) problem as a parameterized mean-field game (PMFG), aiming to reduce complexity via an infinite-population limit. We first show that when dynamics and rewards are Lipschitz, the finite-N ID objective is approximated by the PMFG at rate O( 1 / √ N ). Moreover, beyond the Lipschitz-continuous setting, we prove the same O( 1 / √ N ) decay for the important special case of sequential auctions, despite discontinuities in dynamics, through a tailored auction-specific analysis. Built on our novel approximation results, we further introduce our Adjoint Mean-Field Incentive Design (AMID) algorithm, which uses explicit differentiation of iterated equilibrium operators to compute gradients efficiently. By uniting approximation bounds with optimization guarantees, AMID delivers a powerful, scalable algorithmic tool for many-agent (large N ) ID. Across diverse auction settings, the proposed AMID method substantially increases revenue over first-price formats and outperforms existing benchmark methods.

(entropy-regularized) sum of rewards of player i ∈ [N ] is defined as

we call π π π * a Nash equilibrium (DG-NE) with respect to parameter θ. The set of all Nash equilibria for θ ∈ Θ is denoted Nash τ G (θ). One is typically interested in maximizing a function of the aggregated population behavior (e.g., revenue, negative congestion):

Theorem 1 mirrors bounds in MFGs without ID [75] and Stackelberg MFGs in other settings [54].

For clarity, the theorem is stated as an approximation result for a PDG G that exactly satisfies agent exchangeability, which might not always be the case. In some applications, finding the mean-field formulation M given a PDF G might be nontrivial. The converse problem of constructing an appropriate PMFG M that approximates a given G was studied in [76], and this work can be trivially generalized to this case with an additional approximation bias due to asymmetries in G. The case of auction design, which also does not satisfy Assumption 1 and the assumption of P h,θ being a product measure, will require specific treatment in Section 3.

Having satisfied (D1) for ID with Lipschitz dynamics in the previous section, we turn to (D2)formulating algorithmic methods to solve (MID). We state the standard definitions of value and q-functions for PMFGs, which will be important for learning NEs:

We define the commonly used online mirror descent update rule F omd : Θ × Π H → Π H with

for some given learning rate η > 0 and entropy regularization τ ≥ 0. F omd has received particular attention in MFG literature due to its theoretical and empirical properties. Abbreviating F (T ) omd (θ, ζ) := F omd (θ, F omd (θ, . . . F omd (θ, ζ) . . .)), i.e., F omd (θ, •) applied T times, the repeated iterations F (T ) omd for T > 0 are known to convert to NE for monotone MFGs theoretically [56,79,37], and empirically find good approximations to NE [19] for general MFGs. Furthermore, any π * ∈ Nash τ M (θ) is guaranteed to be a fixed point of the map F omd (θ, •) for some learning rate η. We formulate an explicit differentiation scheme for the PMFG using these properties of F omd . Defining the softmax transform softmax : R [H]×S×A → Π as softmax(ζ)(h, s, a) := expζ h,s,a a ′ expζ h,s,a ′ , the above OMD update rule can be reformulated in terms of log probabilities:

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext 74aa654d-cf64-4869-acbd-e8bae4215a26

引用它的顶会 Paper1

问问它们各自怎么用它

它引用的顶会 Paper18

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖