Towards Deeper Deep Reinforcement Learning with Spectral Normalization
Johan Bjorck, Carla P. Gomes, Kilian Q. Weinberger
摘要
In computer vision and natural language processing, innovations in model architecture that lead to increases in model capacity have reliably translated into gains in performance. In stark contrast with this trend, state-of-the-art reinforcement learning (RL) algorithms often use only small MLPs, and gains in performance typically originate from algorithmic innovations. It is natural to hypothesize that small datasets in RL necessitate simple models to avoid overfitting; however, this hypothesis is untested. In this paper we investigate how RL agents are affected by exchanging the small MLPs with larger modern networks with skip connections and normalization, focusing specifically on soft actor-critic (SAC) algorithms. We verify, empirically, that naively adopting such architectures leads to instabilities and poor performance, likely contributing to the popularity of simple models in practice. However, we show that dataset size is not the limiting factor, and instead argue that intrinsic instability from the actor in SAC taking gradients through the critic is the culprit. We demonstrate that a simple smoothing method can mitigate this issue, which enables stable training with large modern architectures. After smoothing, larger models yield dramatic performance improvements for state-of-the-art agents -- suggesting that more "easy" gains may be had by focusing on model architectures in addition to algorithmic innovations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper24
- PLASTIC: Improving Input and Label Plasticity for Sample Efficient Reinforcement LearningHojoon Lee, Hanseul Cho, Hyunseung Kim, Daehoon Gwak 等NeurIPS 2023 · 被引用 50 次
- Overestimation, Overfitting, and Plasticity in Actor-Critic: the Bitter Lesson of Reinforcement LearningMichal Nauman, Michal Bortkiewicz, Piotr Milos, Tomasz Trzcinski 等ICML 2024 · 被引用 46 次
- Is High Variance Unavoidable in RL? A Case Study in Continuous ControlJohan Bjorck, Carla P. Gomes, Kilian Q. WeinbergerICLR 2022 · 被引用 36 次
- In value-based deep reinforcement learning, a pruned network is a good networkJohan S. Obando-Ceron, Aaron C. Courville, Pablo Samuel CastroICML 2024 · 被引用 36 次
- Stable Gradients for Stable Learning at Scale in Deep Reinforcement LearningRoger Creus Castanyer, Johan S. Obando-Ceron, Lu Li, Pierre-Luc Bacon 等NeurIPS 2025 · 被引用 26 次
它引用的顶会 Paper28
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 被引用 1,852 次
- On Layer Normalization in the Transformer ArchitectureRuibin Xiong, Yunchang Yang, Di He, Kai Zheng 等ICML 2020 · 被引用 1,388 次
相关 Paper
- Hyperspherical Normalization for Scalable Deep Reinforcement LearningHojoon Lee, Youngdo Lee, Takuma Seno, Donghu Kim 等ICML 2025
- The Impact of On-Policy Parallelized Data Collection on Deep Reinforcement Learning NetworksWalter Mayor, Johan S. Obando-Ceron, Aaron C. Courville, Pablo Samuel CastroICML 2025
- SimBa: Simplicity Bias for Scaling Up Parameters in Deep Reinforcement LearningHojoon Lee, Dongyoon Hwang, Donghu Kim, Hyunseung Kim 等ICLR 2025
- Co-Adaptation of Algorithmic and Implementational Innovations in Inference-based Deep Reinforcement LearningHiroki Furuta, Tadashi Kozuno, Tatsuya Matsushima, Yutaka Matsuo 等NeurIPS 2021 · 被引用 14 次
- Striving for Simplicity and Performance in Off-Policy DRL: Output Normalization and Non-Uniform SamplingChe Wang, Yanqiu Wu, Quan Vuong, Keith W. RossICML 2020 · 被引用 38 次
