A Novel Framework for Policy Mirror Descent with General Parameterization and Linear Convergence
Carlo Alfano, Rui Yuan, Patrick Rebeschini
Abstract
Modern policy optimization methods in reinforcement learning, such as TRPO and PPO, owe their success to the use of parameterized policies. However, while theoretical guarantees have been established for this class of algorithms, especially in the tabular setting, the use of general parameterization schemes remains mostly unjustified. In this work, we introduce a framework for policy optimization based on mirror descent that naturally accommodates general parameterizations. The policy class induced by our scheme recovers known classes, e.g., softmax, and generates new ones depending on the choice of mirror map. Using our framework, we obtain the first result that guarantees linear convergence for a policy-gradient-based method involving general parameterization. To demonstrate the ability of our framework to accommodate general parameterization schemes, we provide its sample complexity when using shallow neural networks, show that it represents an improvement upon the previous best results, and empirically validate the effectiveness of our theoretical claims on classic control tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 62507e1e-09fc-4f07-8682-2d7062bb2f88Cited by top-tier papers7
- Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-constraintWei Xiong, Hanze Dong, Chenlu Ye, Ziqi Wang et al.ICML 2024 · 346 citations
- Beyond Stationarity: Convergence Analysis of Stochastic Softmax Policy Gradient MethodsSara Klein, Simon Weissmann, Leif DöringICLR 2024 · 12 citations
- Policy Mirror Descent with LookaheadKimon Protopapas, Anas BarakatNeurIPS 2024 · 7 citations
- On the Second-Order Convergence of Biased Policy Gradient AlgorithmsSiqiao Mu, Diego KlabjanICML 2024 · 4 citations
- Distributionally Robust Reinforcement Learning from Human FeedbackDebmalya Mandal, Paulius Sasnauskas, Goran RadanovicICML 2026
Builds on22
- On the Global Convergence Rates of Softmax Policy Gradient MethodsJincheng Mei, Chenjun Xiao, Csaba Szepesvári, Dale SchuurmansICML 2020 · 349 citations
- Neural Policy Gradient Methods: Global Optimality and Rates of ConvergenceLingxiao Wang, Qi Cai, Zhuoran Yang, Zhaoran WangICLR 2020 · 270 citations
- Natural Policy Gradient Primal-Dual Method for Constrained Markov Decision ProcessesDongsheng Ding, Kaiqing Zhang, Tamer Basar, Mihailo R. JovanovicNeurIPS 2020 · 252 citations
- Bilinear Classes: A Structural Framework for Provable Generalization in RLSimon S. Du, Sham M. Kakade, Jason D. Lee, Shachar Lovett et al.ICML 2021 · 207 citations
- Adaptive Trust Region Policy Optimization: Global Convergence and Faster Rates for Regularized MDPsLior Shani, Yonathan Efroni, Shie MannorAAAI 2020 · 201 citations
Related papers
- Mirror Learning: A Unifying Framework of Policy OptimisationJakub Grudzien Kuba, Christian A. Schröder de Witt, Jakob N. FoersterICML 2022 · 41 citations
- PPO-Clip Attains Global Optimality: Towards Deeper Understandings of ClippingNai-Chieh Huang, Ping-Chun Hsieh, Kuo-Hao Ho, I-Chen WuAAAI 2024 · 34 citations
- Mirror Descent Policy OptimizationManan Tomar, Lior Shani, Yonathan Efroni, Mohammad GhavamzadehICLR 2022 · 111 citations
- A Parametric Class of Approximate Gradient Updates for Policy OptimizationRamki Gummadi, Saurabh Kumar, Junfeng Wen, Dale SchuurmansICML 2022
- Convergence of Policy Mirror Descent Beyond Compatible Function ApproximationUri Sherman, Tomer Koren, Yishay MansourICML 2025
