Recurrent neural networks: vanishing and exploding gradients are not the end of the story
Nicolas Zucchet, Antonio Orvieto
Abstract
Recurrent neural networks (RNNs) notoriously struggle to learn long-term memories, primarily due to vanishing and exploding gradients. The recent success of state-space models (SSMs), a subclass of RNNs, to overcome such difficulties challenges our theoretical understanding. In this paper, we delve into the optimization challenges of RNNs and discover that, as the memory of a network increases, changes in its parameters result in increasingly large output variations, making gradient-based learning highly sensitive, even without exploding gradients. Our analysis further reveals the importance of the element-wise recurrence design pattern combined with careful parametrizations in mitigating this effect. This feature is present in SSMs, as well as in other architectures, such as LSTMs. Overall, our insights provide a new explanation for some of the difficulties in gradient-based learning of RNNs and why some architectures perform better than others.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ac72bec2-d78e-4825-b8cb-9d32d2380237Cited by top-tier papers13
- Theoretical Foundations of Deep Selective State-Space ModelsNicola Muca Cirone, Antonio Orvieto, Benjamin Walker, Cristopher Salvi et al.NeurIPS 2024 · 97 citations
- In Search of Adam's Secret SauceAntonio Orvieto, Robert GowerNeurIPS 2025 · 43 citations
- Provable Benefits of Complex Parameterizations for Structured State Space ModelsYuval Ran-Milo, Eden Lumbroso, Edo Cohen-Karlik, Raja Giryes et al.NeurIPS 2024 · 14 citations
- Predictability Enables Parallelization of Nonlinear State Space ModelsXavier Gonzalez, Leo Kozachkov, David M. Zoltowski, Kenneth L. Clarkson et al.NeurIPS 2025 · 12 citations
- Weight-Space Linear Recurrent Neural NetworksRoussel Desmond Nzoyem, Nawid Keshtmand, Enrique Crespo-Fernandez, Idriss Tsayem et al.ICLR 2026 · 6 citations
Builds on13
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- Combining Recurrent, Convolutional, and Continuous-time Models with Linear State Space LayersAlbert Gu, Isys Johnson, Karan Goel, Khaled Saab et al.NeurIPS 2021 · 1,280 citations
- Resurrecting Recurrent Neural Networks for Long SequencesAntonio Orvieto, Samuel L. Smith, Albert Gu, Anushan Fernando et al.ICML 2023 · 474 citations
Related papers
- StableSSM: Alleviating the Curse of Memory in State-space Models through Stable ReparameterizationShida Wang, Qianxiao LiICML 2024 · 25 citations
- UnICORNN: A recurrent model for learning very long time dependenciesT. Konstantin Rusch, Siddhartha MishraICML 2021 · 76 citations
- On the difficulty of learning chaotic dynamics with RNNsJonas M. Mikhaeil, Zahra Monfared, Daniel DurstewitzNeurIPS 2022 · 109 citations
- Implicit Bias of Linear RNNsMelikasadat Emami, Mojtaba Sahraee-Ardakan, Parthe Pandit, Sundeep Rangan et al.ICML 2021 · 14 citations
- Long Expressive Memory for Sequence ModelingT. Konstantin Rusch, Siddhartha Mishra, N. Benjamin Erichson, Michael W. MahoneyICLR 2022 · 57 citations
