Universality of Linear Recurrences Followed by Non-linear Projections: Finite-Width Guarantees and Benefits of Complex Eigenvalues
Antonio Orvieto, Soham De, Caglar Gulcehre, Razvan Pascanu, Samuel L. Smith
摘要
Deep neural networks based on linear RNNs interleaved with position-wise MLPs are gaining traction as competitive approaches for sequence modeling. Examples of such architectures include state-space models (SSMs) like S4, LRU, and Mamba: recently proposed models that achieve promising performance on text, genetics, and other data that require long-range reasoning. Despite experimental evidence highlighting these architectures' effectiveness and computational efficiency, their expressive power remains relatively unexplored, especially in connection to specific choices crucial in practice -e.g., carefully designed initialization distribution and potential use of complex numbers. In this paper, we show that combining MLPs with both real or complex linear diagonal recurrences leads to arbitrarily precise approximation of regular causal sequence-tosequence maps. At the heart of our proof, we rely on a separation of concerns: the linear RNN provides a lossless encoding of the input sequence, and the MLP performs non-linear processing on this encoding. While we show that real diagonal linear recurrences are enough to achieve universality in this architecture, we prove that employing complex eigenvalues near unit disk -i.e., empirically the most successful strategy in S4 -greatly helps the RNN in storing information. We connect this finding with the vanishing gradient issue and provide experiments supporting our claims. Note: The preliminary version of this manuscript (ICML workshop version) only contains a subset of the results. Our follow-up (Cirone et al., 2024) covers expressive power of gated SSMs such as Mamba (Gu & Dao, 2023) .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Recurrent neural networks: vanishing and exploding gradients are not the end of the storyNicolas Zucchet, Antonio OrvietoNeurIPS 2024 · 被引用 78 次
- The Expressive Capacity of State Space Models: A Formal Language PerspectiveYash Raj Sarrof, Yana Veitsman, Michael HahnNeurIPS 2024 · 被引用 53 次
- Understanding the Differences in Foundation Models: Attention, State Space Models, and Recurrent Neural NetworksJerome Sieber, Carmen Amo Alonso, Alexandre Didier, Melanie N. Zeilinger 等NeurIPS 2024 · 被引用 38 次
- Provable Benefits of Complex Parameterizations for Structured State Space ModelsYuval Ran-Milo, Eden Lumbroso, Edo Cohen-Karlik, Raja Giryes 等NeurIPS 2024 · 被引用 14 次
- Selective Rotary Position EmbeddingSajad Movahedi, Timur Carstensen, Arshia Afzal, Frank Hutter 等ICLR 2026 · 被引用 11 次
它引用的顶会 Paper22
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 被引用 2,665 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
相关 Paper
- Resurrecting Recurrent Neural Networks for Long SequencesAntonio Orvieto, Samuel L. Smith, Albert Gu, Anushan Fernando 等ICML 2023 · 被引用 474 次
- Fixed-Point RNNs: Interpolating from Diagonal to DenseSajad Movahedi, Felix Sarnthein, Nicola Muca Cirone, Antonio OrvietoNeurIPS 2025 · 被引用 5 次
- The Illusion of State in State-Space ModelsWilliam Merrill, Jackson Petty, Ashish SabharwalICML 2024 · 被引用 157 次
- Theoretical Foundations of Deep Selective State-Space ModelsNicola Muca Cirone, Antonio Orvieto, Benjamin Walker, Cristopher Salvi 等NeurIPS 2024 · 被引用 97 次
- Parallelization of Non-linear State-Space Models: Scaling Up Liquid-Resistance Liquid-Capacitance Networks for Efficient Sequence ModelingMónika Farsang, Radu GrosuNeurIPS 2025 · 被引用 14 次
