ϕ-Update: A Class of Policy Update Methods with Policy Convergence Guarantee
Wenye Li, Jiacai Liu, Ke Wei
摘要
Inspired by the similar update pattern of softmax natural policy gradient and Hadamard policy gradient, we propose to study a general policy update rule called ϕ-update, where ϕ refers to a scaling function on advantage functions. Under very mild conditions on ϕ, the global asymptotic state value convergence of ϕ-update is firstly established. Then we show that the policy produced by ϕ-update indeed converges, even when there are multiple optimal policies. This is in stark contrast to existing results where explicit regularizations are required to guarantee the convergence of the policy. Since softmax natural policy gradient is an instance of ϕ-update, it provides an affirmative answer to the question whether the policy produced by softmax natural policy gradient converges. The exact asymptotic convergence rate of state values is further established based on the policy convergence. Lastly, we establish the global linear convergence of ϕ-update.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- On the Global Convergence Rates of Softmax Policy Gradient MethodsJincheng Mei, Chenjun Xiao, Csaba Szepesvári, Dale SchuurmansICML 2020 · 被引用 349 次
- Adaptive Trust Region Policy Optimization: Global Convergence and Faster Rates for Regularized MDPsLior Shani, Yonathan Efroni, Shie MannorAAAI 2020 · 被引用 201 次
- Mirror Descent Policy OptimizationManan Tomar, Lior Shani, Yonathan Efroni, Mohammad GhavamzadehICLR 2022 · 被引用 111 次
- Escaping the Gravitational Pull of SoftmaxJincheng Mei, Chenjun Xiao, Bo Dai, Lihong Li 等NeurIPS 2020 · 被引用 56 次
- Leveraging Non-uniformity in First-order Non-convex OptimizationJincheng Mei, Yue Gao, Bo Dai, Csaba Szepesvári 等ICML 2021 · 被引用 55 次
相关 Paper
- Ordering-based Conditions for Global Convergence of Policy Gradient MethodsJincheng Mei, Bo Dai, Alekh Agarwal, Mohammad Ghavamzadeh 等NeurIPS 2023 · 被引用 4 次
- Convergence of Policy Gradient for Entropy Regularized MDPs with Neural Network Approximation in the Mean-Field RegimeJames-Michael Leahy, Bekzhan Kerimkulov, David Siska, Lukasz SzpruchICML 2022 · 被引用 23 次
- On the Global Convergence Rates of Decentralized Softmax Gradient Play in Markov Potential GamesRunyu Zhang, Jincheng Mei, Bo Dai, Dale Schuurmans 等NeurIPS 2022 · 被引用 38 次
- A Parametric Class of Approximate Gradient Updates for Policy OptimizationRamki Gummadi, Saurabh Kumar, Junfeng Wen, Dale SchuurmansICML 2022
- Federated Natural Policy Gradient and Actor Critic Methods for Multi-task Reinforcement LearningTong Yang, Shicong Cen, Yuting Wei, Yuxin Chen 等NeurIPS 2024 · 被引用 14 次
