Calibration of Shared Equilibria in General Sum Partially Observable Markov Games
Nelson Vadori, Sumitra Ganesh, Prashant P. Reddy, Manuela Veloso
Abstract
Training multi-agent systems (MAS) to achieve realistic equilibria gives us a useful tool to understand and model real-world systems. We consider a general sum partially observable Markov game where agents of different types share a single policy network, conditioned on agent-specific information. This paper aims at i) formally understanding equilibria reached by such agents, and ii) matching emergent phenomena of such equilibria to real-world targets. Parameter sharing with decentralized execution has been introduced as an efficient way to train multiple agents using a single policy network. However, the nature of resulting equilibria reached by such agents is not yet understood: we introduce the novel concept of Shared equilibrium as a symmetric pure Nash equilibrium of a certain Functional Form Game (FFG) and prove convergence to the latter for a certain class of games using self-play. In addition, it is important that such equilibria satisfy certain constraints so that MAS are calibrated to real world data for practical use: we solve this problem by introducing a novel dual-Reinforcement Learning based approach that fits emergent behaviors of agents in a Shared equilibrium to externally-specified targets, and apply our methods to a -player market example. We do so by calibrating parameters governing distributions of agent types rather than individual agents, which allows both behavior differentiation among agents and coherent scaling of the shared policy network to multiple agents.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Evolution Strategies for Approximate Solution of Bayesian GamesZun Li, Michael P. WellmanAAAI 2021 · 19 citations
- Learning in Stackelberg Mean Field Games: A Non-Asymptotic AnalysisSihan Zeng, Benjamin Patrick Evans, Sujay Bhatt, Leo Ardon et al.NeurIPS 2025 · 1 citation
Builds on1
Related papers
- Generalization in Mean Field Games by Learning Master PoliciesSarah Perrin, Mathieu Laurière, Julien Pérolat, Romuald Élie et al.AAAI 2022 · 47 citations
- AgentMixer: Multi-Agent Correlated Policy FactorizationZhiyuan Li, Wenshuai Zhao, Lijun Wu, Joni PajarinenAAAI 2025 · 7 citations
- Learning Nash Equilibrium of Markov Potential Games with a Shared Constraint via Primal-Dual OptimizationSongtao Feng, Michael R. Dorothy, Jie FuAAAI 2025
- Sample-Efficient Reinforcement Learning of Partially Observable Markov GamesQinghua Liu, Csaba Szepesvári, Chi JinNeurIPS 2022 · 43 citations
- Scaling Multi-Agent Reinforcement Learning with Selective Parameter SharingFilippos Christianos, Georgios Papoudakis, Arrasy Rahman, Stefano V. AlbrechtICML 2021 · 165 citations
