Learning Low Dimensional State Spaces with Overparameterized Recurrent Neural Nets
Edo Cohen-Karlik, Itamar Menuhin-Gruman, Raja Giryes, Nadav Cohen, Amir Globerson
Abstract
Overparameterization in deep learning typically refers to settings where a trained neural network (NN) has representational capacity to fit the training data in many ways, some of which generalize well, while others do not. In the case of Recurrent Neural Networks (RNNs), there exists an additional layer of overparameterization, in the sense that a model may exhibit many solutions that generalize well for sequence lengths seen in training, some of which extrapolate to longer sequences, while others do not. Numerous works have studied the tendency of Gradient Descent (GD) to fit overparameterized NNs with solutions that generalize well. On the other hand, its tendency to fit overparameterized RNNs with solutions that extrapolate has been discovered only recently and is far less understood. In this paper, we analyze the extrapolation properties of GD when applied to overparameterized linear RNNs. In contrast to recent arguments suggesting an implicit bias towards short-term memory, we provide theoretical evidence for learning low-dimensional state spaces, which can also model long-term memory. Our result relies on a dynamical characterization which shows that GD (with small step size and near-zero initialization) strives to maintain a certain form of balancedness, as well as on tools developed in the context of the moment problem from statistics (recovery of a probability distribution from its moments). Experiments corroborate our theory, demonstrating extrapolation via learning low-dimensional state spaces with both linear and non-linear RNNs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e0520622-75a5-477b-96e2-0b41a52baf2aCited by top-tier papers8
- Provable Benefits of Complex Parameterizations for Structured State Space ModelsYuval Ran-Milo, Eden Lumbroso, Edo Cohen-Karlik, Raja Giryes et al.NeurIPS 2024 · 14 citations
- Why is Your Language Model a Poor Implicit Reward Model?Noam Razin, Yong Lin, Jiarui Yao, Sanjeev AroraICLR 2026 · 8 citations
- PAC-Bayes Generalisation Bounds for Dynamical Systems including Stable RNNsDeividas Eringis, John Leth, Zheng-Hua Tan, Rafael Wisniewski et al.AAAI 2024 · 5 citations
- Learning Dynamics of RNNs in Closed-Loop EnvironmentsYoav Ger, Omri BarakNeurIPS 2025 · 2 citations
- The Implicit Bias of Structured State Space Models Can Be Poisoned With Clean LabelsYonatan Slutzky, Yotam Alexander, Noam Razin, Nadav CohenNeurIPS 2025 · 2 citations
Builds on16
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- Combining Recurrent, Convolutional, and Continuous-time Models with Linear State Space LayersAlbert Gu, Isys Johnson, Karan Goel, Khaled Saab et al.NeurIPS 2021 · 1,280 citations
- Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationOfir Press, Noah A. Smith, Mike LewisICLR 2022 · 1,168 citations
- HiPPO: Recurrent Memory with Optimal Polynomial ProjectionsAlbert Gu, Tri Dao, Stefano Ermon, Atri Rudra et al.NeurIPS 2020 · 1,100 citations
- Diagonal State Spaces are as Effective as Structured State SpacesAnkit Gupta, Albert Gu, Jonathan BerantNeurIPS 2022 · 546 citations
Related papers
- Inverse Approximation Theory for Nonlinear Recurrent Neural NetworksShida Wang, Zhong Li, Qianxiao LiICLR 2024 · 10 citations
- Algorithm Development in Neural Networks: Insights from the Streaming Parity TaskLoek van Rossem, Andrew M. SaxeICML 2025
- On the Curse of Memory in Recurrent Neural Networks: Approximation and Optimization AnalysisZhong Li, Jiequn Han, Weinan E, Qianxiao LiICLR 2021 · 40 citations
- Implicit Bias of Linear RNNsMelikasadat Emami, Mojtaba Sahraee-Ardakan, Parthe Pandit, Sundeep Rangan et al.ICML 2021 · 14 citations
- MomentumRNN: Integrating Momentum into Recurrent Neural NetworksTan M. Nguyen, Richard G. Baraniuk, Andrea L. Bertozzi, Stanley J. Osher et al.NeurIPS 2020 · 32 citations
