Improving Deep Policy Gradients with Value Function Search
Enrico Marchesini, Christopher Amato
Abstract
Deep Policy Gradient (PG) algorithms employ value networks to drive the learning of parameterized policies and reduce the variance of the gradient estimates. However, value function approximation gets stuck in local optima and struggles to fit the actual return, limiting the variance reduction efficacy and leading policies to sub-optimal performance. This paper focuses on improving value approximation and analyzing the effects on Deep PG primitives such as value prediction, variance reduction, and correlation of gradient estimates with the true gradient. To this end, we introduce a Value Function Search that employs a population of perturbed value networks to search for a better approximation. Our framework does not require additional environment interactions, gradient computations, or ensembles, providing a computationally inexpensive approach to enhance the supervised learning task on which value networks train. Crucially, we show that improving Deep PG primitives results in improved sample efficiency and policies with higher returns using common continuous control benchmark domains.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f9fcb97f-0ff7-4c7c-9b4b-2e1aaf99912aCited by top-tier papers2
- EvoRainbow: Combining Improvements in Evolutionary Reinforcement Learning for Policy SearchPengyi Li, Yan Zheng, Hongyao Tang, Xian Fu et al.ICML 2024 · 13 citations
- LaRes: Evolutionary Reinforcement Learning with LLM-based Adaptive Reward SearchPengyi Li, Hongyao Tang, Jinbin Qiao, Yan Zheng et al.NeurIPS 2025 · 7 citations
Builds on7
- QPLEX: Duplex Dueling Multi-Agent Q-LearningJianhao Wang, Zhizhou Ren, Terry Liu, Yang Yu et al.ICLR 2021 · 595 citations
- SUNRISE: A Simple Unified Framework for Ensemble Learning in Deep Reinforcement LearningKimin Lee, Michael Laskin, Aravind Srinivas, Pieter AbbeelICML 2021 · 239 citations
- Understanding and Preventing Capacity Loss in Reinforcement LearningClare Lyle, Mark Rowland, Will DabneyICLR 2022 · 151 citations
- A Closer Look at Deep Policy GradientsAndrew Ilyas, Logan Engstrom, Shibani Santurkar, Dimitris Tsipras et al.ICLR 2020 · 107 citations
- Proximal Distilled Evolutionary Reinforcement LearningCristian Bodnar, Ben Day, Pietro LióAAAI 2020 · 101 citations
Related papers
- Deep Bayesian Quadrature Policy OptimizationRavi Tej Akella, Kamyar Azizzadenesheli, Mohammad Ghavamzadeh, Animashree Anandkumar et al.AAAI 2021 · 5 citations
- Deterministic Value-Policy GradientsQingpeng Cai, Ling Pan, Pingzhong TangAAAI 2020 · 1 citation
- An Improved Analysis of (Variance-Reduced) Policy Gradient and Natural Policy Gradient MethodsYanli Liu, Kaiqing Zhang, Tamer Basar, Wotao YinNeurIPS 2020 · 128 citations
- PAGE-PG: A Simple and Loopless Variance-Reduced Policy Gradient Method with Probabilistic Gradient EstimationMatilde Gargiani, Andrea Zanelli, Andrea Martinelli, Tyler H. Summers et al.ICML 2022 · 17 citations
- From Importance Sampling to Doubly Robust Policy GradientJiawei Huang, Nan JiangICML 2020 · 26 citations
