Scalable Neural Incentive Design with Parameterized Mean-Field Approximation
Nathan Corecco, Batuhan Yardim, Vinzenz Thoma, Zebang Shen, Niao He
摘要
Designing incentives for a multi-agent system to induce a desirable Nash equilibrium is both a crucial and challenging problem appearing in many decision-making domains, especially for a large number of agents N . Under the exchangeability assumption, we formalize this incentive design (ID) problem as a parameterized mean-field game (PMFG), aiming to reduce complexity via an infinite-population limit. We first show that when dynamics and rewards are Lipschitz, the finite-N ID objective is approximated by the PMFG at rate O( 1 / √ N ). Moreover, beyond the Lipschitz-continuous setting, we prove the same O( 1 / √ N ) decay for the important special case of sequential auctions, despite discontinuities in dynamics, through a tailored auction-specific analysis. Built on our novel approximation results, we further introduce our Adjoint Mean-Field Incentive Design (AMID) algorithm, which uses explicit differentiation of iterated equilibrium operators to compute gradients efficiently. By uniting approximation bounds with optimization guarantees, AMID delivers a powerful, scalable algorithmic tool for many-agent (large N ) ID. Across diverse auction settings, the proposed AMID method substantially increases revenue over first-price formats and outperforms existing benchmark methods.
(entropy-regularized) sum of rewards of player i ∈ [N ] is defined as
we call π π π * a Nash equilibrium (DG-NE) with respect to parameter θ. The set of all Nash equilibria for θ ∈ Θ is denoted Nash τ G (θ). One is typically interested in maximizing a function of the aggregated population behavior (e.g., revenue, negative congestion):
Theorem 1 mirrors bounds in MFGs without ID [75] and Stackelberg MFGs in other settings [54].
For clarity, the theorem is stated as an approximation result for a PDG G that exactly satisfies agent exchangeability, which might not always be the case. In some applications, finding the mean-field formulation M given a PDF G might be nontrivial. The converse problem of constructing an appropriate PMFG M that approximates a given G was studied in [76], and this work can be trivially generalized to this case with an additional approximation bias due to asymmetries in G. The case of auction design, which also does not satisfy Assumption 1 and the assumption of P h,θ being a product measure, will require specific treatment in Section 3.
Having satisfied (D1) for ID with Lipschitz dynamics in the previous section, we turn to (D2)formulating algorithmic methods to solve (MID). We state the standard definitions of value and q-functions for PMFGs, which will be important for learning NEs:
We define the commonly used online mirror descent update rule F omd : Θ × Π H → Π H with
for some given learning rate η > 0 and entropy regularization τ ≥ 0. F omd has received particular attention in MFG literature due to its theoretical and empirical properties. Abbreviating F (T ) omd (θ, ζ) := F omd (θ, F omd (θ, . . . F omd (θ, ζ) . . .)), i.e., F omd (θ, •) applied T times, the repeated iterations F (T ) omd for T > 0 are known to convert to NE for monotone MFGs theoretically [56,79,37], and empirically find good approximations to NE [19] for general MFGs. Furthermore, any π * ∈ Nash τ M (θ) is guaranteed to be a fixed point of the map F omd (θ, •) for some learning rate η. We formulate an explicit differentiation scheme for the PMFG using these properties of F omd . Defining the softmax transform softmax : R [H]×S×A → Π as softmax(ζ)(h, s, a) := expζ h,s,a a ′ expζ h,s,a ′ , the above OMD update rule can be reformulated in terms of log probabilities:
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper18
- Fictitious Play for Mean Field Games: Continuous Time Analysis and ApplicationsSarah Perrin, Julien Pérolat, Mathieu Laurière, Matthieu Geist 等NeurIPS 2020 · 被引用 150 次
- Implicit Learning Dynamics in Stackelberg Games: Equilibria Characterization, Convergence Analysis, and Empirical StudyTanner Fiez, Benjamin Chasnov, Lillian J. RatliffICML 2020 · 被引用 144 次
- End-to-End Learning and Intervention in GamesJiayang Li, Jing Yu, Yu Marco Nie, Zhaoran WangNeurIPS 2020 · 被引用 48 次
- Generalization in Mean Field Games by Learning Master PoliciesSarah Perrin, Mathieu Laurière, Julien Pérolat, Romuald Élie 等AAAI 2022 · 被引用 47 次
- Policy Mirror Ascent for Efficient and Independent Learning in Mean Field GamesBatuhan Yardim, Semih Cayci, Matthieu Geist, Niao HeICML 2023 · 被引用 33 次
相关 Paper
- Actor-Critic Provably Finds Nash Equilibria of Linear-Quadratic Mean-Field GamesZuyue Fu, Zhuoran Yang, Yongxin Chen, Zhaoran WangICLR 2020 · 被引用 61 次
- Last Iterate Convergence in Monotone Mean Field GamesNoboru Isobe, Kenshi Abe, Kaito AriuNeurIPS 2025 · 被引用 2 次
- Learning While Playing in Mean-Field Games: Convergence and OptimalityQiaomin Xie, Zhuoran Yang, Zhaoran Wang, Andreea MincaICML 2021 · 被引用 45 次
- Policy Gradient Methods Converge Globally in Imperfect-Information Extensive-Form GamesFivos Kalogiannis, Gabriele FarinaNeurIPS 2025 · 被引用 2 次
- On the Convergence of Model Free Learning in Mean Field GamesRomuald Elie, Julien Pérolat, Mathieu Laurière, Matthieu Geist 等AAAI 2020 · 被引用 101 次
