Policy Optimization with Stochastic Mirror Descent
Long Yang, Yu Zhang, Gang Zheng, Qian Zheng, Pengfei Li, Jianhang Huang, Gang Pan
Abstract
Improving sample efficiency has been a longstanding goal in reinforcement learning. This paper proposes VRMPO algorithm: a sample efficient policy gradient method with stochastic mirror descent. In VRMPO, a novel variance-reduced policy gradient estimator is presented to improve sample efficiency. We prove that the proposed VRMPO needs only O( -3 ) sample trajectories to achieve an -approximate first-order stationary point, which matches the best sample complexity for policy optimization. The extensive experimental results demonstrate that VRMPO outperforms the state-of-the-art policy gradient methods in various settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 17cc077d-938e-4897-89e2-d8b199e56ec8Cited by top-tier papers3
- Sample Efficient Policy Gradient Methods with Recursive Variance ReductionPan Xu, Felicia Gao, Quanquan GuICLR 2020 · 99 citations
- On the Hidden Biases of Policy Mirror Ascent in Continuous Action SpacesAmrit Singh Bedi, Souradip Chakraborty, Anjaly Parayil, Brian M. Sadler et al.ICML 2022 · 20 citations
- Controlling Type Confounding in Ad Hoc Teamwork with Instance-wise Teammate Feedback RectificationDong Xing, Pengjie Gu, Qian Zheng, Xinrun Wang et al.ICML 2023 · 4 citations
Builds on5
- Adaptive Trust Region Policy Optimization: Global Convergence and Faster Rates for Regularized MDPsLior Shani, Yonathan Efroni, Shie MannorAAAI 2020 · 201 citations
- Mirror Descent Policy OptimizationManan Tomar, Lior Shani, Yonathan Efroni, Mohammad GhavamzadehICLR 2022 · 111 citations
- Sample Efficient Policy Gradient Methods with Recursive Variance ReductionPan Xu, Felicia Gao, Quanquan GuICLR 2020 · 99 citations
- Sample Complexity of Policy Gradient Finding Second-Order Stationary PointsLong Yang, Qian Zheng, Gang PanAAAI 2021 · 25 citations
- Bregman Gradient Policy OptimizationFeihu Huang, Shangqian Gao, Heng HuangICLR 2022 · 19 citations
Related papers
- Momentum-Based Policy Gradient MethodsFeihu Huang, Shangqian Gao, Jian Pei, Heng HuangICML 2020 · 47 citations
- Momentum-Based Federated Reinforcement Learning with Interaction and Communication EfficiencySheng Yue, Xingyuan Hua, Lili Chen, Ju RenINFOCOM 2024 · 4 citations
- PAGE-PG: A Simple and Loopless Variance-Reduced Policy Gradient Method with Probabilistic Gradient EstimationMatilde Gargiani, Andrea Zanelli, Andrea Martinelli, Tyler H. Summers et al.ICML 2022 · 17 citations
- Reflective Policy OptimizationYaozhong Gan, Renye Yan, Zhe Wu, Junliang XingICML 2024
- Greedy-GQ with Variance Reduction: Finite-time Analysis and Improved ComplexityShaocong Ma, Ziyi Chen, Yi Zhou, Shaofeng ZouICLR 2021 · 12 citations
