On Proximal Policy Optimization's Heavy-tailed Gradients
Saurabh Garg, Joshua Zhanson, Emilio Parisotto, Adarsh Prasad, J. Zico Kolter, Zachary C. Lipton, Sivaraman Balakrishnan, Ruslan Salakhutdinov, Pradeep Ravikumar
Abstract
Modern policy gradient algorithms such as Proximal Policy Optimization (PPO) rely on an arsenal of heuristics, including loss clipping and gradient clipping, to ensure successful learning. These heuristics are reminiscent of techniques from robust statistics, commonly used for estimation in outlier-rich (``heavy-tailed'') regimes. In this paper, we present a detailed empirical study to characterize the heavy-tailed nature of the gradients of the PPO surrogate reward function. We demonstrate that the gradients, especially for the actor network, exhibit pronounced heavy-tailedness and that it increases as the agent's policy diverges from the behavioral policy (i.e., as the agent goes further off policy). Further examination implicates the likelihood ratios and advantages in the surrogate reward as the main sources of the observed heavy-tailedness. We then highlight issues arising due to the heavy-tailed nature of the gradients. In this light, we study the effects of the standard PPO clipping heuristics, demonstrating that these tricks primarily serve to offset heavy-tailedness in gradients. Thus motivated, we propose incorporating GMOM, a high-dimensional robust estimator, into PPO as a substitute for three clipping tricks. Despite requiring less hyperparameter tuning, our method matches the performance of PPO (with all heuristics enabled) on a battery of MuJoCo continuous control tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e1c53f6c-889b-49ca-ac02-e65e3b9b2353Cited by top-tier papers11
- Towards Robust Offline Reinforcement Learning under Diverse Data CorruptionRui Yang, Han Zhong, Jiawei Xu, Amy Zhang et al.ICLR 2024 · 28 citations
- Eliminating Sharp Minima from SGD with Truncated Heavy-tailed NoiseXingyu Wang, Sewoong Oh, Chang-Han RheeICLR 2022 · 21 citations
- On the Hidden Biases of Policy Mirror Ascent in Continuous Action SpacesAmrit Singh Bedi, Souradip Chakraborty, Anjaly Parayil, Brian M. Sadler et al.ICML 2022 · 20 citations
- Clipped Gradient Methods for Nonsmooth Convex Optimization under Heavy-Tailed Noise: A Refined AnalysisZijian LiuICLR 2026 · 5 citations
- Robust Policy Expansion for Offline-to-Online RL under Diverse Data CorruptionLongxiang He, Deheng Ye, Junbo Tan, Xueqian Wang et al.NeurIPS 2025 · 5 citations
Builds on6
- Why Gradient Clipping Accelerates Training: A Theoretical Justification for AdaptivityJingzhao Zhang, Tianxing He, Suvrit Sra, Ali JadbabaieICLR 2020 · 598 citations
- Implementation Matters in Deep RL: A Case Study on PPO and TRPOLogan Engstrom, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras et al.ICLR 2020 · 305 citations
- A Diffusion Theory For Deep Learning Dynamics: Stochastic Gradient Descent Exponentially Favors Flat MinimaZeke Xie, Issei Sato, Masashi SugiyamaICLR 2021 · 165 citations
- A Closer Look at Deep Policy GradientsAndrew Ilyas, Logan Engstrom, Shibani Santurkar, Dimitris Tsipras et al.ICLR 2020 · 107 citations
- Fractional Underdamped Langevin Dynamics: Retargeting SGD with Momentum under Heavy-Tailed Gradient NoiseUmut Simsekli, Lingjiong Zhu, Yee Whye Teh, Mert GürbüzbalabanICML 2020 · 58 citations
Related papers
- Provably Robust Temporal Difference Learning for Heavy-Tailed RewardsSemih Cayci, Atilla EryilmazNeurIPS 2023 · 12 citations
- Exact Policy Recovery in Offline RL with Both Heavy-Tailed Rewards and Data CorruptionYiding Chen, Xuezhou Zhang, Qiaomin Xie, Xiaojin ZhuAAAI 2024 · 2 citations
- Truncated Gaussian Policy for Debiased Continuous ControlGanghun Lee, Minji Kim, Minsu Lee, Byoung-Tak ZhangAAAI 2025 · 1 citation
- The Sufficiency of Off-Policyness and Soft Clipping: PPO Is Still Insufficient according to an Off-Policy MeasureXing Chen, Dongcui Diao, Hechang Chen, Hengshuai Yao et al.AAAI 2023 · 28 citations
- Off-Policy Proximal Policy OptimizationWenjia Meng, Qian Zheng, Gang Pan, Yilong YinAAAI 2023 · 27 citations
