Stable Gradients for Stable Learning at Scale in Deep Reinforcement Learning
Roger Creus Castanyer, Johan S. Obando-Ceron, Lu Li, Pierre-Luc Bacon, Glen Berseth, Aaron C. Courville, Pablo Samuel Castro
Abstract
Scaling deep reinforcement learning networks is challenging and often results in degraded performance, yet the root causes of this failure mode remain poorly understood. Several recent works have proposed mechanisms to address this, but they are often complex and fail to highlight the causes underlying this difficulty. In this work, we conduct a series of empirical analyses which suggest that the combination of non-stationarity with gradient pathologies, due to suboptimal architectural choices, underlie the challenges of scale. We propose a series of direct interventions that stabilize gradient flow, enabling robust performance across a range of network depths and widths. Our interventions are simple to implement and compatible with well-established algorithms, and result in an effective mechanism that enables strong performance even at large scales. We validate our findings on a variety of agents and suites of environments. Source code here.
"We must be able to look at the world and see it as a dynamic process, not a static picture."
-David Bohm
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- XQC: Well-conditioned Optimization Accelerates Deep Reinforcement LearningDaniel Palenicek, Florian Vogt, Joe Watson, Ingmar Posner et al.ICLR 2026 · 20 citations
- Measure gradients, not activations! Enhancing neuronal activity in deep reinforcement learningJiashun Liu, Zihao Wu, Johan S. Obando-Ceron, Pablo Samuel Castro et al.NeurIPS 2025 · 15 citations
- Simplicial Embeddings Improve Sample Efficiency in Actor–Critic AgentsJohan Obando-Ceron, Walter Mayor, Samuel Lavoie, Scott Fujimoto et al.ICLR 2026 · 12 citations
- TQL: Scaling Q-Functions with Transformers by Preventing Attention CollapsePerry Dong, Kuo-Han Hung, Alexander Swerdlow, Dorsa Sadigh et al.ICML 2026 · 7 citations
- Stabilizing Policy Gradients for Sample-Efficient Reinforcement Learning in LLM ReasoningLuckeciano Carvalho Melo, Alessandro Abate, Yarin GalICLR 2026 · 7 citations
Builds on35
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville et al.NeurIPS 2021 · 1,067 citations
- AdaBelief Optimizer: Adapting Stepsizes by the Belief in Observed GradientsJuntang Zhuang, Tommy Tang, Yifan Ding, Sekhar Tatikonda et al.NeurIPS 2020 · 697 citations
- A Signal Propagation Perspective for Pruning Neural Networks at InitializationNamhoon Lee, Thalaiyasingam Ajanthan, Stephen Gould, Philip H. S. TorrICLR 2020 · 174 citations
- Understanding Plasticity in Neural NetworksClare Lyle, Zeyu Zheng, Evgenii Nikishin, Bernardo Ávila Pires et al.ICML 2023 · 162 citations
- Implicit Under-Parameterization Inhibits Data-Efficient Deep Reinforcement LearningAviral Kumar, Rishabh Agarwal, Dibya Ghosh, Sergey LevineICLR 2021 · 155 citations
Related papers
- Mind the GAP! The Challenges of Scale in Pixel-based Deep Reinforcement LearningGhada Sokar, Pablo Samuel CastroNeurIPS 2025 · 5 citations
- Network Sparsity Unlocks the Scaling Potential of Deep Reinforcement LearningGuozheng Ma, Lu Li, Zilin Wang, Li Shen et al.ICML 2025
- Understanding and Preventing Capacity Loss in Reinforcement LearningClare Lyle, Mark Rowland, Will DabneyICLR 2022 · 151 citations
- Scalable Option Learning in High-Throughput EnvironmentsMikael Henaff, Scott Fujimoto, Michael Matthews, Michael RabbatICML 2026 · 5 citations
- ScaleMoE: Mixture-of-Experts for Scalable Continuous Control in Actor-Critic Reinforcement LearningYi Ma, Chenjun Xiao, Hongyao Tang, Yaodong Yang et al.ICML 2026
