Global optimality of softmax policy gradient with single hidden layer neural networks in the mean-field regime
Andrea Agazzi, Jianfeng Lu
Abstract
We study the problem of policy optimization for infinite-horizon discounted Markov Decision with softmax policy and nonlinear function approximation trained with policy algorithms. We concentrate on the training dynamics in the mean-field regime, e.g., the behavior of wide single hidden layer neural networks, when exploration encouraged through entropy regularization. The dynamics of these models is established a Wasserstein gradient flow of distributions in parameter space. We further prove global of the fixed points of this dynamics under mild conditions on their initialization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2ee8ce1e-87ee-4ca3-88a8-02ad9b204db0Cited by top-tier papers10
- Policy Optimization for Continuous Reinforcement LearningHanyang Zhao, Wenpin Tang, David D. YaoNeurIPS 2023 · 47 citations
- A multiscale analysis of mean-field transformers in the moderate interaction regimeGiuseppe Bruno, Federico Pasqualotto, Andrea AgazziNeurIPS 2025 · 29 citations
- Convergence of Policy Gradient for Entropy Regularized MDPs with Neural Network Approximation in the Mean-Field RegimeJames-Michael Leahy, Bekzhan Kerimkulov, David Siska, Lukasz SzpruchICML 2022 · 23 citations
- Wasserstein Flow Meets Replicator Dynamics: A Mean-Field Analysis of Representation Learning in Actor-CriticYufeng Zhang, Siyu Chen, Zhuoran Yang, Michael I. Jordan et al.NeurIPS 2021 · 6 citations
- Limiting fluctuation and trajectorial stability of multilayer neural networks with mean field trainingHuy Tuan Pham, Phan-Minh NguyenNeurIPS 2021 · 6 citations
Builds on4
- On the Global Convergence Rates of Softmax Policy Gradient MethodsJincheng Mei, Chenjun Xiao, Csaba Szepesvári, Dale SchuurmansICML 2020 · 349 citations
- Neural Policy Gradient Methods: Global Optimality and Rates of ConvergenceLingxiao Wang, Qi Cai, Zhuoran Yang, Zhaoran WangICLR 2020 · 270 citations
- PC-PG: Policy Cover Directed Exploration for Provable Policy Gradient LearningAlekh Agarwal, Mikael Henaff, Sham M. Kakade, Wen SunNeurIPS 2020 · 126 citations
- Can Temporal-Difference and Q-Learning Learn Representation? A Mean-Field TheoryYufeng Zhang, Qi Cai, Zhuoran Yang, Yongxin Chen et al.NeurIPS 2020 · 12 citations
Related papers
- Global optimality of Elman-type RNNs in the mean-field regimeAndrea Agazzi, Jianfeng Lu, Sayan MukherjeeICML 2023 · 2 citations
- Beyond Stationarity: Convergence Analysis of Stochastic Softmax Policy Gradient MethodsSara Klein, Simon Weissmann, Leif DöringICLR 2024 · 12 citations
- Mirror Mean-Field Langevin DynamicsAnming Gu, Juno KimICML 2026 · 3 citations
- Structure Matters: Dynamic Policy GradientSara Klein, Xiangyuan Zhang, Tamer Basar, Simon Weissmann et al.NeurIPS 2025 · 1 citation
- Improved statistical and computational complexity of the mean-field Langevin dynamics under structured dataAtsushi Nitanda, Kazusato Oko, Taiji Suzuki, Denny WuICLR 2024 · 4 citations
