Model-Based Reparameterization Policy Gradient Methods: Theory and Practical Algorithms
Shenao Zhang, Boyi Liu, Zhaoran Wang, Tuo Zhao
Abstract
ReParameterization (RP) Policy Gradient Methods (PGMs) have been widely adopted for continuous control tasks in robotics and computer graphics. However, recent studies have revealed that, when applied to long-term reinforcement learning problems, model-based RP PGMs may experience chaotic and non-smooth optimization landscapes with exploding gradient variance, which leads to slow convergence. This is in contrast to the conventional belief that reparameterization methods have low gradient estimation variance in problems such as training deep generative models. To comprehend this phenomenon, we conduct a theoretical examination of model-based RP PGMs and search for solutions to the optimization difficulties. Specifically, we analyze the convergence of the model-based RP PGMs and pinpoint the smoothness of function approximators as a major factor that affects the quality of gradient estimation. Based on our analysis, we propose a spectral normalization method to mitigate the exploding variance issue caused by long model unrolls. Our experimental results demonstrate that proper normalization significantly reduces the gradient variance of model-based RP PGMs. As a result, the performance of the proposed method is comparable or superior to other gradient estimators, such as the Likelihood Ratio (LR) gradient estimator. Our code is available at https://github.com/agentification/RP_PGM .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Do Transformer World Models Give Better Policy Gradients?Michel Ma, Tianwei Ni, Clement Gehring, Pierluca D'Oro et al.ICML 2024 · 7 citations
- Adaptive-Gradient Policy Optimization: Enhancing Policy Learning in Non-Smooth Differentiable SimulationsFeng Gao, Liangzhi Shi, Shenao Zhang, Zhaoran Wang et al.ICML 2024 · 7 citations
- Learning to Reason as Action Abstractions with Scalable Mid-Training RLShenao Zhang, Donghan Yu, Yihao Feng, Bowen Jin et al.ICLR 2026 · 5 citations
- Distributions as Actions: A Unified Framework for Diverse Action SpacesJiamin He, A. Rupam Mahmood, Martha WhiteICLR 2026
- Reparameterization Proximal Policy OptimizationHai Zhong, Xun Wang, Zhuoran Li, Longbo HuangICML 2026
Builds on11
- Model Based Reinforcement Learning for AtariLukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski et al.ICLR 2020 · 969 citations
- Neural Policy Gradient Methods: Global Optimality and Rates of ConvergenceLingxiao Wang, Qi Cai, Zhuoran Yang, Zhaoran WangICLR 2020 · 270 citations
- Accelerated Policy Learning with Parallel Differentiable SimulationJie Xu, Viktor Makoviychuk, Yashraj Narang, Fabio Ramos et al.ICLR 2022 · 141 citations
- Sample Efficient Policy Gradient Methods with Recursive Variance ReductionPan Xu, Felicia Gao, Quanquan GuICLR 2020 · 99 citations
- Deep Reinforcement Learning with Robust and Smooth PolicyQianli Shen, Yan Li, Haoming Jiang, Zhaoran Wang et al.ICML 2020 · 95 citations
Related papers
- Fractal Landscapes in Policy OptimizationTao Wang, Sylvia L. Herbert, Sicun GaoNeurIPS 2023 · 10 citations
- Reparameterized Policy Learning for Multimodal Trajectory OptimizationZhiao Huang, Litian Liang, Zhan Ling, Xuanlin Li et al.ICML 2023 · 21 citations
- Reparameterization Flow Policy OptimizationHai Zhong, Zhuoran Li, Xun Wang, Longbo HuangICML 2026
- Non-Uniform Noise-to-Signal Ratio in the REINFORCE Policy-Gradient EstimatorHaoyu Han, Heng YangICML 2026 · 3 citations
- Mollification Effects of Policy Gradient MethodsTao Wang, Sylvia L. Herbert, Sicun GaoICML 2024 · 2 citations
