Lune

ICML2023顶会

Policy Mirror Ascent for Efficient and Independent Learning in Mean Field Games

Batuhan Yardim, Semih Cayci, Matthieu Geist, Niao He

2023年份
33被引次数
11顶会引用

摘要

Mean-field games have been used as a theoretical tool to obtain an approximate Nash equilibrium for symmetric and anonymous NN-player games. However, limiting applicability, existing theoretical results assume variations of a"population generative model", which allows arbitrary modifications of the population distribution by the learning algorithm. Moreover, learning algorithms typically work on abstract simulators with population instead of the NN-player game. Instead, we show that NN agents running policy mirror ascent converge to the Nash equilibrium of the regularized game within O~(ε−2)\widetilde{\mathcal{O}}(\varepsilon^{-2}) samples from a single sample trajectory without a population generative model, up to a standard O(1N)\mathcal{O}(\frac{1}{\sqrt{N}}) error due to the mean field. Taking a divergent approach from the literature, instead of working with the best-response map we first show that a policy mirror ascent map can be used to construct a contractive operator having the Nash equilibrium as its fixed point. We analyze single-path TD learning for NN-agent games, proving sample complexity guarantees by only using a sample path from the NN-agent simulator without a population generative model. Furthermore, we demonstrate that our methodology allows for independent learning by NN agents with finite sample guarantees.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper11

问问它们各自怎么用它

它引用的顶会 Paper11

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖