Inverse Approximation Theory for Nonlinear Recurrent Neural Networks
Shida Wang, Zhong Li, Qianxiao Li
摘要
We prove an inverse approximation theorem for the approximation of nonlinear sequence-to-sequence relationships using recurrent neural networks (RNNs). This is a so-called Bernstein-type result in approximation theory, which deduces properties of a target function under the assumption that it can be effectively approximated by a hypothesis space. In particular, we show that nonlinear sequence relationships that can be stably approximated by nonlinear RNNs must have an exponential decaying memory structure - a notion that can be made precise. This extends the previously identified curse of memory in linear RNNs into the general nonlinear setting, and quantifies the essential limitations of the RNN architecture for learning sequential relationships with long-term memory. Based on the analysis, we propose a principled reparameterization method to overcome the limitations. Our theoretical results are confirmed by numerical experiments. The code has been released in https://github.com/radarFudan/Curse-of-memory
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Recurrent neural networks: vanishing and exploding gradients are not the end of the storyNicolas Zucchet, Antonio OrvietoNeurIPS 2024 · 被引用 78 次
- State-space models with layer-wise nonlinearity are universal approximators with exponential decaying memoryShida Wang, Beichen XueNeurIPS 2023 · 被引用 49 次
- StableSSM: Alleviating the Curse of Memory in State-space Models through Stable ReparameterizationShida Wang, Qianxiao LiICML 2024 · 被引用 25 次
- From Generalization Analysis to Optimization Designs for State Space ModelsFusheng Liu, Qianxiao LiICML 2024 · 被引用 12 次
它引用的顶会 Paper5
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
- HiPPO: Recurrent Memory with Optimal Polynomial ProjectionsAlbert Gu, Tri Dao, Stefano Ermon, Atri Rudra 等NeurIPS 2020 · 被引用 1,100 次
- Simplified State Space Layers for Sequence ModelingJimmy T. H. Smith, Andrew Warrington, Scott W. LindermanICLR 2023 · 被引用 78 次
- State-space models with layer-wise nonlinearity are universal approximators with exponential decaying memoryShida Wang, Beichen XueNeurIPS 2023 · 被引用 49 次
- On the Curse of Memory in Recurrent Neural Networks: Approximation and Optimization AnalysisZhong Li, Jiequn Han, Weinan E, Qianxiao LiICLR 2021 · 被引用 40 次
相关 Paper
- On the approximation properties of recurrent encoder-decoder architecturesZhong Li, Haotian Jiang, Qianxiao LiICLR 2022 · 被引用 8 次
- Learning Low Dimensional State Spaces with Overparameterized Recurrent Neural NetsEdo Cohen-Karlik, Itamar Menuhin-Gruman, Raja Giryes, Nadav Cohen 等ICLR 2023 · 被引用 2 次
- Identifying nonlinear dynamical systems with multiple time scales and long-range dependenciesDominik Schmidt, Georgia Koppe, Zahra Monfared, Max Beutelspacher 等ICLR 2021 · 被引用 41 次
- UnICORNN: A recurrent model for learning very long time dependenciesT. Konstantin Rusch, Siddhartha MishraICML 2021 · 被引用 76 次
- Implicit Bias of Linear RNNsMelikasadat Emami, Mojtaba Sahraee-Ardakan, Parthe Pandit, Sundeep Rangan 等ICML 2021 · 被引用 14 次
