Stable Gradients for Stable Learning at Scale in Deep Reinforcement Learning
Roger Creus Castanyer, Johan S. Obando-Ceron, Lu Li, Pierre-Luc Bacon, Glen Berseth, Aaron C. Courville, Pablo Samuel Castro
摘要
Scaling deep reinforcement learning networks is challenging and often results in degraded performance, yet the root causes of this failure mode remain poorly understood. Several recent works have proposed mechanisms to address this, but they are often complex and fail to highlight the causes underlying this difficulty. In this work, we conduct a series of empirical analyses which suggest that the combination of non-stationarity with gradient pathologies, due to suboptimal architectural choices, underlie the challenges of scale. We propose a series of direct interventions that stabilize gradient flow, enabling robust performance across a range of network depths and widths. Our interventions are simple to implement and compatible with well-established algorithms, and result in an effective mechanism that enables strong performance even at large scales. We validate our findings on a variety of agents and suites of environments. Source code here.
"We must be able to look at the world and see it as a dynamic process, not a static picture."
-David Bohm
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- XQC: Well-conditioned Optimization Accelerates Deep Reinforcement LearningDaniel Palenicek, Florian Vogt, Joe Watson, Ingmar Posner 等ICLR 2026 · 被引用 20 次
- Measure gradients, not activations! Enhancing neuronal activity in deep reinforcement learningJiashun Liu, Zihao Wu, Johan S. Obando-Ceron, Pablo Samuel Castro 等NeurIPS 2025 · 被引用 15 次
- Simplicial Embeddings Improve Sample Efficiency in Actor–Critic AgentsJohan Obando-Ceron, Walter Mayor, Samuel Lavoie, Scott Fujimoto 等ICLR 2026 · 被引用 12 次
- TQL: Scaling Q-Functions with Transformers by Preventing Attention CollapsePerry Dong, Kuo-Han Hung, Alexander Swerdlow, Dorsa Sadigh 等ICML 2026 · 被引用 7 次
- Stabilizing Policy Gradients for Sample-Efficient Reinforcement Learning in LLM ReasoningLuckeciano Carvalho Melo, Alessandro Abate, Yarin GalICLR 2026 · 被引用 7 次
它引用的顶会 Paper35
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville 等NeurIPS 2021 · 被引用 1,067 次
- AdaBelief Optimizer: Adapting Stepsizes by the Belief in Observed GradientsJuntang Zhuang, Tommy Tang, Yifan Ding, Sekhar Tatikonda 等NeurIPS 2020 · 被引用 697 次
- A Signal Propagation Perspective for Pruning Neural Networks at InitializationNamhoon Lee, Thalaiyasingam Ajanthan, Stephen Gould, Philip H. S. TorrICLR 2020 · 被引用 174 次
- Understanding Plasticity in Neural NetworksClare Lyle, Zeyu Zheng, Evgenii Nikishin, Bernardo Ávila Pires 等ICML 2023 · 被引用 162 次
- Implicit Under-Parameterization Inhibits Data-Efficient Deep Reinforcement LearningAviral Kumar, Rishabh Agarwal, Dibya Ghosh, Sergey LevineICLR 2021 · 被引用 155 次
相关 Paper
- Mind the GAP! The Challenges of Scale in Pixel-based Deep Reinforcement LearningGhada Sokar, Pablo Samuel CastroNeurIPS 2025 · 被引用 5 次
- Network Sparsity Unlocks the Scaling Potential of Deep Reinforcement LearningGuozheng Ma, Lu Li, Zilin Wang, Li Shen 等ICML 2025
- Understanding and Preventing Capacity Loss in Reinforcement LearningClare Lyle, Mark Rowland, Will DabneyICLR 2022 · 被引用 151 次
- Scalable Option Learning in High-Throughput EnvironmentsMikael Henaff, Scott Fujimoto, Michael Matthews, Michael RabbatICML 2026 · 被引用 5 次
- ScaleMoE: Mixture-of-Experts for Scalable Continuous Control in Actor-Critic Reinforcement LearningYi Ma, Chenjun Xiao, Hongyao Tang, Yaodong Yang 等ICML 2026
