A Parametric Class of Approximate Gradient Updates for Policy Optimization
Ramki Gummadi, Saurabh Kumar, Junfeng Wen, Dale Schuurmans
摘要
Approaches to policy optimization have been motivated from diverse principles, based on how the parametric model is interpreted (e.g. value versus policy representation) or how the learning objective is formulated, yet they share a common goal of maximizing expected return. To better capture the commonalities and identify key differences between policy optimization methods, we develop a unified perspective that re-expresses the underlying updates in terms of a limited choice of gradient form and scaling function. In particular, we identify a parameterized space of approximate gradient updates for policy optimization that is highly structured, yet covers both classical and recent examples, including PPO. As a result, we obtain novel yet well motivated updates that generalize existing algorithms in a way that can deliver benefits both in terms of convergence speed and final result quality. An experimental investigation demonstrates that the additional degrees of freedom provided in the parameterized family of updates can be leveraged to obtain non-trivial improvements both in synthetic domains and on popular deep RL benchmarks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Leverage the Average: an Analysis of KL Regularization in Reinforcement LearningNino Vieillard, Tadashi Kozuno, Bruno Scherrer, Olivier Pietquin 等NeurIPS 2020 · 被引用 106 次
- An operator view of policy gradient methodsDibya Ghosh, Marlos C. Machado, Nicolas Le RouxNeurIPS 2020 · 被引用 30 次
- Characterizing the Gap Between Actor-Critic and Policy GradientJunfeng Wen, Saurabh Kumar, Ramki Gummadi, Dale SchuurmansICML 2021 · 被引用 18 次
- Learning Value Functions in Deep Policy Gradients using Residual VarianceYannis Flet-Berliac, Reda Ouhamma, Odalric-Ambrym Maillard, Philippe PreuxICLR 2021 · 被引用 16 次
相关 Paper
- Mirror Learning: A Unifying Framework of Policy OptimisationJakub Grudzien Kuba, Christian A. Schröder de Witt, Jakob N. FoersterICML 2022 · 被引用 41 次
- A Novel Framework for Policy Mirror Descent with General Parameterization and Linear ConvergenceCarlo Alfano, Rui Yuan, Patrick RebeschiniNeurIPS 2023 · 被引用 25 次
- Taylor Expansion of Discount FactorsYunhao Tang, Mark Rowland, Rémi Munos, Michal ValkoICML 2021 · 被引用 8 次
- Distributional Meta-Gradient Reinforcement LearningHaiyan Yin, Shuicheng Yan, Zhongwen XuICLR 2023
- A Closer Look at Deep Policy GradientsAndrew Ilyas, Logan Engstrom, Shibani Santurkar, Dimitris Tsipras 等ICLR 2020 · 被引用 107 次
