Implicit Bias of Linear RNNs
Melikasadat Emami, Mojtaba Sahraee-Ardakan, Parthe Pandit, Sundeep Rangan, Alyson K. Fletcher
Abstract
Contemporary wisdom based on empirical studies suggests that standard recurrent neural networks (RNNs) do not perform well on tasks requiring long-term memory. However, precise reasoning for this behavior is still unknown. This paper provides a rigorous explanation of this property in the special case of linear RNNs. Although this work is limited to linear RNNs, even these systems have traditionally been difficult to analyze due to their non-linear parameterization. Using recently-developed kernel regime analysis, our main result shows that linear RNNs learned from random initializations are functionally equivalent to a certain weighted 1D-convolutional network. Importantly, the weightings in the equivalent model cause an implicit bias to elements with smaller time lags in the convolution and hence, shorter memory. The degree of this bias depends on the variance of the transition kernel matrix at initialization and is related to the classic exploding and vanishing gradients problem. The theory is validated in both synthetic and real data experiments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7a75301d-308b-4316-90be-a72d4b8443caCited by top-tier papers9
- Noisy Recurrent Neural NetworksSoon Hoe Lim, N. Benjamin Erichson, Liam Hodgkinson, Michael W. MahoneyNeurIPS 2021 · 77 citations
- How connectivity structure shapes rich and lazy learning in neural circuitsYuhan Helena Liu, Aristide Baratin, Jonathan Cornford, Stefan Mihalas et al.ICLR 2024 · 26 citations
- PAC-Bayes Generalisation Bounds for Dynamical Systems including Stable RNNsDeividas Eringis, John Leth, Zheng-Hua Tan, Rafael Wisniewski et al.AAAI 2024 · 5 citations
- Separation and Bias of Deep Equilibrium Models on Expressivity and Learning DynamicsZhoutong Wu, Yimu Zhang, Cong Fang, Zhouchen LinNeurIPS 2024 · 3 citations
- The Implicit Bias of Structured State Space Models Can Be Poisoned With Clean LabelsYonatan Slutzky, Yotam Alexander, Noam Razin, Nadav CohenNeurIPS 2025 · 2 citations
Builds on3
- RNNPool: Efficient Non-linear Pooling for RAM Constrained InferenceOindrila Saha, Aditya Kusupati, Harsha Vardhan Simhadri, Manik Varma et al.NeurIPS 2020 · 58 citations
- Generalization Error of Generalized Linear Models in High DimensionsMelikasadat Emami, Mojtaba Sahraee-Ardakan, Parthe Pandit, Sundeep Rangan et al.ICML 2020 · 40 citations
- The Recurrent Neural Tangent KernelSina Alemohammad, Zichao Wang, Randall Balestriero, Richard G. BaraniukICLR 2021 · 6 citations
Related papers
- Revisiting Glorot Initialization for Long-Range Linear RecurrencesNoga Bar, Mariia Seleznova, Yotam Alexander, Gitta Kutyniok et al.NeurIPS 2025 · 3 citations
- Identifying nonlinear dynamical systems with multiple time scales and long-range dependenciesDominik Schmidt, Georgia Koppe, Zahra Monfared, Max Beutelspacher et al.ICLR 2021 · 41 citations
- Recurrent neural networks: vanishing and exploding gradients are not the end of the storyNicolas Zucchet, Antonio OrvietoNeurIPS 2024 · 78 citations
- On the difficulty of learning chaotic dynamics with RNNsJonas M. Mikhaeil, Zahra Monfared, Daniel DurstewitzNeurIPS 2022 · 109 citations
- On the Curse of Memory in Recurrent Neural Networks: Approximation and Optimization AnalysisZhong Li, Jiequn Han, Weinan E, Qianxiao LiICLR 2021 · 40 citations
