Global optimality of softmax policy gradient with single hidden layer neural networks in the mean-field regime
Andrea Agazzi, Jianfeng Lu
2021年份
3被引次数
10顶会引用
摘要
We study the problem of policy optimization for infinite-horizon discounted Markov Decision with softmax policy and nonlinear function approximation trained with policy algorithms. We concentrate on the training dynamics in the mean-field regime, e.g., the behavior of wide single hidden layer neural networks, when exploration encouraged through entropy regularization. The dynamics of these models is established a Wasserstein gradient flow of distributions in parameter space. We further prove global of the fixed points of this dynamics under mild conditions on their initialization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Policy Optimization for Continuous Reinforcement LearningHanyang Zhao, Wenpin Tang, David D. YaoNeurIPS 2023 · 被引用 47 次
- A multiscale analysis of mean-field transformers in the moderate interaction regimeGiuseppe Bruno, Federico Pasqualotto, Andrea AgazziNeurIPS 2025 · 被引用 29 次
- Convergence of Policy Gradient for Entropy Regularized MDPs with Neural Network Approximation in the Mean-Field RegimeJames-Michael Leahy, Bekzhan Kerimkulov, David Siska, Lukasz SzpruchICML 2022 · 被引用 23 次
- Wasserstein Flow Meets Replicator Dynamics: A Mean-Field Analysis of Representation Learning in Actor-CriticYufeng Zhang, Siyu Chen, Zhuoran Yang, Michael I. Jordan 等NeurIPS 2021 · 被引用 6 次
- Limiting fluctuation and trajectorial stability of multilayer neural networks with mean field trainingHuy Tuan Pham, Phan-Minh NguyenNeurIPS 2021 · 被引用 6 次
它引用的顶会 Paper4
- On the Global Convergence Rates of Softmax Policy Gradient MethodsJincheng Mei, Chenjun Xiao, Csaba Szepesvári, Dale SchuurmansICML 2020 · 被引用 349 次
- Neural Policy Gradient Methods: Global Optimality and Rates of ConvergenceLingxiao Wang, Qi Cai, Zhuoran Yang, Zhaoran WangICLR 2020 · 被引用 270 次
- PC-PG: Policy Cover Directed Exploration for Provable Policy Gradient LearningAlekh Agarwal, Mikael Henaff, Sham M. Kakade, Wen SunNeurIPS 2020 · 被引用 126 次
- Can Temporal-Difference and Q-Learning Learn Representation? A Mean-Field TheoryYufeng Zhang, Qi Cai, Zhuoran Yang, Yongxin Chen 等NeurIPS 2020 · 被引用 12 次
相关 Paper
- Global optimality of Elman-type RNNs in the mean-field regimeAndrea Agazzi, Jianfeng Lu, Sayan MukherjeeICML 2023 · 被引用 2 次
- Beyond Stationarity: Convergence Analysis of Stochastic Softmax Policy Gradient MethodsSara Klein, Simon Weissmann, Leif DöringICLR 2024 · 被引用 12 次
- Mirror Mean-Field Langevin DynamicsAnming Gu, Juno KimICML 2026 · 被引用 3 次
- Structure Matters: Dynamic Policy GradientSara Klein, Xiangyuan Zhang, Tamer Basar, Simon Weissmann 等NeurIPS 2025 · 被引用 1 次
- Improved statistical and computational complexity of the mean-field Langevin dynamics under structured dataAtsushi Nitanda, Kazusato Oko, Taiji Suzuki, Denny WuICLR 2024 · 被引用 4 次
