Momentum-Based Policy Gradient Methods
Feihu Huang, Shangqian Gao, Jian Pei, Heng Huang
摘要
In the paper, we propose a class of efficient momentum-based policy gradient methods for the model-free reinforcement learning, which use adaptive learning rates and do not require any large batches. Specifically, we propose a fast important-sampling momentum-based policy gradient (IS-MBPG) method based on a new momentum-based variance reduced technique and the importance sampling technique. We also propose a fast Hessian-aided momentum-based policy gradient (HA-MBPG) method based on the momentum-based variance reduced technique and the Hessian-aided technique. Moreover, we prove that both the IS-MBPG and HA-MBPG methods reach the best known sample complexity of for finding an -stationary point of the non-concave performance function, which only require one trajectory at each iteration. In particular, we present a non-adaptive version of IS-MBPG method, i.e., IS-MBPG*, which also reaches the best known sample complexity of without any large batches. In the experiments, we apply four benchmark tasks to demonstrate the effectiveness of our algorithms.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- An Improved Analysis of (Variance-Reduced) Policy Gradient and Natural Policy Gradient MethodsYanli Liu, Kaiqing Zhang, Tamer Basar, Wotao YinNeurIPS 2020 · 被引用 128 次
- Accelerated Stochastic Gradient-free and Projection-free MethodsFeihu Huang, Lue Tao, Songcan ChenICML 2020 · 被引用 27 次
- Stochastic Second-Order Methods Improve Best-Known Sample Complexity of SGD for Gradient-Dominated FunctionsSaeed Masiha, Saber Salehkaleybar, Niao He, Negar Kiyavash 等NeurIPS 2022 · 被引用 22 次
- Bregman Gradient Policy OptimizationFeihu Huang, Shangqian Gao, Heng HuangICLR 2022 · 被引用 19 次
- Reinforcement Learning with General Utilities: Simpler Variance Reduction and Large State-Action SpaceAnas Barakat, Ilyas Fatkhullin, Niao HeICML 2023 · 被引用 18 次
它引用的顶会 Paper3
- Neural Policy Gradient Methods: Global Optimality and Rates of ConvergenceLingxiao Wang, Qi Cai, Zhuoran Yang, Zhaoran WangICLR 2020 · 被引用 270 次
- Sample Efficient Policy Gradient Methods with Recursive Variance ReductionPan Xu, Felicia Gao, Quanquan GuICLR 2020 · 被引用 99 次
- Non-asymptotic Convergence of Adam-type Reinforcement Learning Algorithms under Markovian SamplingHuaqing Xiong, Tengyu Xu, Yingbin Liang, Wei ZhangAAAI 2021 · 被引用 37 次
相关 Paper
- MDPGT: Momentum-Based Decentralized Policy Gradient TrackingZhanhong Jiang, Xian Yeow Lee, Sin Yong Tan, Kai Liang Tan 等AAAI 2022 · 被引用 11 次
- Policy Optimization with Stochastic Mirror DescentLong Yang, Yu Zhang, Gang Zheng, Qian Zheng 等AAAI 2022 · 被引用 38 次
- PAGE-PG: A Simple and Loopless Variance-Reduced Policy Gradient Method with Probabilistic Gradient EstimationMatilde Gargiani, Andrea Zanelli, Andrea Martinelli, Tyler H. Summers 等ICML 2022 · 被引用 17 次
- Better SGD using Second-order MomentumHoang Tran, Ashok CutkoskyNeurIPS 2022 · 被引用 18 次
- Stochastic Policy Gradient Methods: Improved Sample Complexity for Fisher-non-degenerate PoliciesIlyas Fatkhullin, Anas Barakat, Anastasia Kireeva, Niao HeICML 2023 · 被引用 61 次
