Lune

ICLR2021Top-tier venue

Global optimality of softmax policy gradient with single hidden layer neural networks in the mean-field regime

Andrea Agazzi, Jianfeng Lu

2021Year
3Citations
10Top-tier citations

Abstract

We study the problem of policy optimization for infinite-horizon discounted Markov Decision with softmax policy and nonlinear function approximation trained with policy algorithms. We concentrate on the training dynamics in the mean-field regime, e.g., the behavior of wide single hidden layer neural networks, when exploration encouraged through entropy regularization. The dynamics of these models is established a Wasserstein gradient flow of distributions in parameter space. We further prove global of the fixed points of this dynamics under mild conditions on their initialization.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 2ee8ce1e-87ee-4ca3-88a8-02ad9b204db0

Cited by top-tier papers10

Ask how each one uses it

Builds on4

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines