Stabilizing Policy Gradient Methods via Reward Profiling
Shihab Ahmed, El Houcine Bergou, Yue Wang, Aritra Dutta
Abstract
Policy gradient methods, which have been extensively studied in the last decade, offer an effective and efficient framework for reinforcement learning problems. However, their performances can often be unsatisfactory, suffering from unreliable reward improvements and slow convergence, due to high variance in gradient estimations. In this paper, we propose a universal reward profiling framework that can be seamlessly integrated with any policy gradient algorithm, where we selectively update the policy based on high-confidence performance estimations. We theoretically justify that our technique will not slow down the convergence of the baseline policy gradient methods, but with high probability, will result in stable and monotonic improvements of their performance. Empirically, on eight continuous‐control benchmarks (Box2D and MuJoCo/PyBullet), our profiling yields up to 1.5x faster convergence to near‐optimal returns, up to 1.75x reduction in return variance on some setups. Our profiling approach offers a general, theoretically grounded path to more reliable and efficient policy learning in complex environments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dbe697f0-00f0-47bd-a059-0bb1f708471dBuilds on12
- Neural Policy Gradient Methods: Global Optimality and Rates of ConvergenceLingxiao Wang, Qi Cai, Zhuoran Yang, Zhaoran WangICLR 2020 · 270 citations
- Closing the Gap: Tighter Analysis of Alternating Stochastic Gradient Methods for Bilevel ProblemsTianyi Chen, Yuejiao Sun, Wotao YinNeurIPS 2021 · 176 citations
- Improving Sample Complexity Bounds for (Natural) Actor-Critic AlgorithmsTengyu Xu, Zhe Wang, Yingbin LiangNeurIPS 2020 · 110 citations
- A Closer Look at Deep Policy GradientsAndrew Ilyas, Logan Engstrom, Shibani Santurkar, Dimitris Tsipras et al.ICLR 2020 · 107 citations
- Provably Convergent Two-Timescale Off-Policy Actor-Critic with Function ApproximationShangtong Zhang, Bo Liu, Hengshuai Yao, Shimon WhitesonICML 2020 · 58 citations
Related papers
- Policy Search by Target Distribution Learning for Continuous ControlChuheng Zhang, Yuanqi Li, Jian LiAAAI 2020 · 6 citations
- Variance Penalized On-Policy and Off-Policy Actor-CriticArushi Jain, Gandharv Patil, Ayush Jain, Khimya Khetarpal et al.AAAI 2021 · 11 citations
- Gradient Information Matters in Policy Optimization by Back-propagating through ModelChongchong Li, Yue Wang, Wei Chen, Yuting Liu et al.ICLR 2022 · 10 citations
- Addressing Action Oscillations through Learning Policy InertiaChen Chen, Hongyao Tang, Jianye Hao, Wulong Liu et al.AAAI 2021 · 27 citations
- Globally Optimal Policy Gradient Algorithms for Reinforcement Learning with PID Control PoliciesVipul Sharma, Wesley Suttle, S. SivaranjaniNeurIPS 2025 · 3 citations
