Normalization and effective learning rates in reinforcement learning
Clare Lyle, Zeyu Zheng, Khimya Khetarpal, James Martens, Hado Philip van Hasselt, Razvan Pascanu, Will Dabney
Abstract
Normalization layers have recently experienced a renaissance in the deep reinforcement learning and continual learning literature, with several works highlighting diverse benefits such as improving loss landscape conditioning and combatting overestimation bias. However, normalization brings with it a subtle but important side effect: an equivalence between growth in the norm of the network parameters and decay in the effective learning rate. This becomes problematic in continual learning settings, where the resulting effective learning rate schedule may decay to near zero too quickly relative to the timescale of the learning problem. We propose to make the learning rate schedule explicit with a simple re-parameterization which we call Normalize-and-Project (NaP), which couples the insertion of normalization layers with weight projection, ensuring that the effective learning rate remains constant throughout training. This technique reveals itself as a powerful analytical tool to better understand learning rate schedules in deep reinforcement learning, and as a means of improving robustness to nonstationarity in synthetic plasticity loss benchmarks along with both the single-task and sequential variants of the Arcade Learning Environment. We also show that our approach can be easily applied to popular architectures such as ResNets and transformers while recovering and in some cases even slightly improving the performance of the base model in common stationary benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 09be303d-9e13-4199-8102-4134ba41846eCited by top-tier papers15
- Stable Gradients for Stable Learning at Scale in Deep Reinforcement LearningRoger Creus Castanyer, Johan S. Obando-Ceron, Lu Li, Pierre-Luc Bacon et al.NeurIPS 2025 · 26 citations
- Scaling Off-Policy Reinforcement Learning with Batch and Weight NormalizationDaniel Palenicek, Florian Vogt, Joe Watson, Jan PetersNeurIPS 2025 · 22 citations
- XQC: Well-conditioned Optimization Accelerates Deep Reinforcement LearningDaniel Palenicek, Florian Vogt, Joe Watson, Ingmar Posner et al.ICLR 2026 · 20 citations
- Weight Decay may matter more than µP for Learning Rate Transfer in PracticeAtli Kosson, Jeremy Welborn, Yang Liu, Martin Jaggi et al.ICLR 2026 · 11 citations
- Stable Deep Reinforcement Learning via Isotropic Gaussian RepresentationsAli Saheb pasand, Johan Obando-Ceron, Aaron Courville, Pouya Bashivan et al.ICML 2026 · 5 citations
Builds on31
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- On Layer Normalization in the Transformer ArchitectureRuibin Xiong, Yunchang Yang, Di He, Kai Zheng et al.ICML 2020 · 1,388 citations
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu et al.ICLR 2020 · 1,170 citations
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville et al.NeurIPS 2021 · 1,067 citations
- Efficient Online Reinforcement Learning with Offline DataPhilip J. Ball, Laura Smith, Ilya Kostrikov, Sergey LevineICML 2023 · 326 citations
Related papers
- Understanding Plasticity in Neural NetworksClare Lyle, Zeyu Zheng, Evgenii Nikishin, Bernardo Ávila Pires et al.ICML 2023 · 162 citations
- Self-Normalized Resets for Plasticity in Continual LearningVivek F. Farias, Adam Daniel JozefiakICLR 2025
- Mitigating Plasticity Loss through Architectural Design in Continual LearningNiklas Koeppe, Luiz Felipe Vecchietti, Dongqi Han, Dongsheng Li et al.ICML 2026
- Natural continual learning: success is a journey, not (just) a destinationTa-Chu Kao, Kristopher T. Jensen, Gido van de Ven, Alberto Bernacchia et al.NeurIPS 2021 · 72 citations
- Activation by Interval-wise Dropout: A Simple Way to Prevent Neural Networks from Plasticity LossSangyeon Park, Isaac Han, Seungwon Oh, Kyung-Joong KimICML 2025
