ICML2026

Recurrent Structural Policy Gradient for Partially Observable Mean Field Games

Clarisse Wibault, Sebastian Towers, Tiphaine Wibault, Juan Duque, Johannes Forkel, George Whittle, Andreas Schaab, Chiyuan Wang, Yucheng Yang, Michael A Osborne, Benjamin Moll, Jakob Foerster

被引用 2 次

摘要

Mean Field Games (MFGs) provide a principled framework for modelling interactions in large population systems. However, algorithmic progress has been limited since model-free methods are high variance and exact methods scale poorly. Recent Hybrid Structural Methods (HSMs) reduce variance while maintaining tractability by leveraging low-dimensional individual state and action spaces and known transition dynamics to compute the exact expected return conditioned on Monte Carlo rollouts of common noise. However, HSMs have not been extended to partially observable settings. We propose Recurrent Structural Policy Gradient (RSPG), the first history-aware HSM for MFGs with public partial information. RSPG achieves an order-of-magnitude faster convergence than model-free RL methods while learning historyaware behaviour, unlike current HSMs. To facilitate research into MFGs, we also introduce MFAX, our JAX-based framework for MFGs that supports both analytic and sample-based meanfield updates. MFAX and usage examples can be found at https://clarisse-wibault. github.io/rspg/ . Introduction Training policies in large multi-agent systems is notoriously difficult: Multi-Agent Reinforcement Learning (MARL) based methods rely on high-variance, trajectory-based sampling and therefore scale poorly as the number of agents † Equal supervision