A Parametric Class of Approximate Gradient Updates for Policy Optimization
Ramki Gummadi, Saurabh Kumar, Junfeng Wen, Dale Schuurmans
Abstract
Approaches to policy optimization have been motivated from diverse principles, based on how the parametric model is interpreted (e.g. value versus policy representation) or how the learning objective is formulated, yet they share a common goal of maximizing expected return. To better capture the commonalities and identify key differences between policy optimization methods, we develop a unified perspective that re-expresses the underlying updates in terms of a limited choice of gradient form and scaling function. In particular, we identify a parameterized space of approximate gradient updates for policy optimization that is highly structured, yet covers both classical and recent examples, including PPO. As a result, we obtain novel yet well motivated updates that generalize existing algorithms in a way that can deliver benefits both in terms of convergence speed and final result quality. An experimental investigation demonstrates that the additional degrees of freedom provided in the parameterized family of updates can be leveraged to obtain non-trivial improvements both in synthetic domains and on popular deep RL benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ed77cb62-ef5b-49ab-a492-e67964cb1f5bBuilds on5
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Leverage the Average: an Analysis of KL Regularization in Reinforcement LearningNino Vieillard, Tadashi Kozuno, Bruno Scherrer, Olivier Pietquin et al.NeurIPS 2020 · 106 citations
- An operator view of policy gradient methodsDibya Ghosh, Marlos C. Machado, Nicolas Le RouxNeurIPS 2020 · 30 citations
- Characterizing the Gap Between Actor-Critic and Policy GradientJunfeng Wen, Saurabh Kumar, Ramki Gummadi, Dale SchuurmansICML 2021 · 18 citations
- Learning Value Functions in Deep Policy Gradients using Residual VarianceYannis Flet-Berliac, Reda Ouhamma, Odalric-Ambrym Maillard, Philippe PreuxICLR 2021 · 16 citations
Related papers
- Mirror Learning: A Unifying Framework of Policy OptimisationJakub Grudzien Kuba, Christian A. Schröder de Witt, Jakob N. FoersterICML 2022 · 41 citations
- A Novel Framework for Policy Mirror Descent with General Parameterization and Linear ConvergenceCarlo Alfano, Rui Yuan, Patrick RebeschiniNeurIPS 2023 · 25 citations
- Taylor Expansion of Discount FactorsYunhao Tang, Mark Rowland, Rémi Munos, Michal ValkoICML 2021 · 8 citations
- Distributional Meta-Gradient Reinforcement LearningHaiyan Yin, Shuicheng Yan, Zhongwen XuICLR 2023
- A Closer Look at Deep Policy GradientsAndrew Ilyas, Logan Engstrom, Shibani Santurkar, Dimitris Tsipras et al.ICLR 2020 · 107 citations
