Model-Based Reparameterization Policy Gradient Methods: Theory and Practical Algorithms
Shenao Zhang, Boyi Liu, Zhaoran Wang, Tuo Zhao
摘要
ReParameterization (RP) Policy Gradient Methods (PGMs) have been widely adopted for continuous control tasks in robotics and computer graphics. However, recent studies have revealed that, when applied to long-term reinforcement learning problems, model-based RP PGMs may experience chaotic and non-smooth optimization landscapes with exploding gradient variance, which leads to slow convergence. This is in contrast to the conventional belief that reparameterization methods have low gradient estimation variance in problems such as training deep generative models. To comprehend this phenomenon, we conduct a theoretical examination of model-based RP PGMs and search for solutions to the optimization difficulties. Specifically, we analyze the convergence of the model-based RP PGMs and pinpoint the smoothness of function approximators as a major factor that affects the quality of gradient estimation. Based on our analysis, we propose a spectral normalization method to mitigate the exploding variance issue caused by long model unrolls. Our experimental results demonstrate that proper normalization significantly reduces the gradient variance of model-based RP PGMs. As a result, the performance of the proposed method is comparable or superior to other gradient estimators, such as the Likelihood Ratio (LR) gradient estimator. Our code is available at https://github.com/agentification/RP_PGM .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Do Transformer World Models Give Better Policy Gradients?Michel Ma, Tianwei Ni, Clement Gehring, Pierluca D'Oro 等ICML 2024 · 被引用 7 次
- Adaptive-Gradient Policy Optimization: Enhancing Policy Learning in Non-Smooth Differentiable SimulationsFeng Gao, Liangzhi Shi, Shenao Zhang, Zhaoran Wang 等ICML 2024 · 被引用 7 次
- Learning to Reason as Action Abstractions with Scalable Mid-Training RLShenao Zhang, Donghan Yu, Yihao Feng, Bowen Jin 等ICLR 2026 · 被引用 5 次
- Distributions as Actions: A Unified Framework for Diverse Action SpacesJiamin He, A. Rupam Mahmood, Martha WhiteICLR 2026
- Reparameterization Proximal Policy OptimizationHai Zhong, Xun Wang, Zhuoran Li, Longbo HuangICML 2026
它引用的顶会 Paper11
- Model Based Reinforcement Learning for AtariLukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski 等ICLR 2020 · 被引用 969 次
- Neural Policy Gradient Methods: Global Optimality and Rates of ConvergenceLingxiao Wang, Qi Cai, Zhuoran Yang, Zhaoran WangICLR 2020 · 被引用 270 次
- Accelerated Policy Learning with Parallel Differentiable SimulationJie Xu, Viktor Makoviychuk, Yashraj Narang, Fabio Ramos 等ICLR 2022 · 被引用 141 次
- Sample Efficient Policy Gradient Methods with Recursive Variance ReductionPan Xu, Felicia Gao, Quanquan GuICLR 2020 · 被引用 99 次
- Deep Reinforcement Learning with Robust and Smooth PolicyQianli Shen, Yan Li, Haoming Jiang, Zhaoran Wang 等ICML 2020 · 被引用 95 次
相关 Paper
- Fractal Landscapes in Policy OptimizationTao Wang, Sylvia L. Herbert, Sicun GaoNeurIPS 2023 · 被引用 10 次
- Reparameterized Policy Learning for Multimodal Trajectory OptimizationZhiao Huang, Litian Liang, Zhan Ling, Xuanlin Li 等ICML 2023 · 被引用 21 次
- Reparameterization Flow Policy OptimizationHai Zhong, Zhuoran Li, Xun Wang, Longbo HuangICML 2026
- Non-Uniform Noise-to-Signal Ratio in the REINFORCE Policy-Gradient EstimatorHaoyu Han, Heng YangICML 2026 · 被引用 3 次
- Mollification Effects of Policy Gradient MethodsTao Wang, Sylvia L. Herbert, Sicun GaoICML 2024 · 被引用 2 次
