On the Curse of Memory in Recurrent Neural Networks: Approximation and Optimization Analysis
Zhong Li, Jiequn Han, Weinan E, Qianxiao Li
摘要
We study the approximation properties and optimization dynamics of recurrent neural networks (RNNs) when applied to learn input-output relationships in temporal data. We consider the simple but representative setting of using continuous-time linear RNNs to learn from data generated by linear relationships. Mathematically, the latter can be understood as a sequence of linear functionals. We prove a universal approximation theorem of such linear functionals, and characterize the approximation rate and its relation with memory. Moreover, we perform a fine-grained dynamical analysis of training linear RNNs, which further reveal the intricate interactions between memory and learning. A unifying theme uncovered is the non-trivial effect of memory, a notion that can be made precise in our framework, on approximation and optimization: when there is long term memory in the target, it takes a large number of neurons to approximate it. Moreover, the training process will suffer from slow downs. In particular, both of these effects become exponentially more pronounced with memory - a phenomenon we call the "curse of memory". These analyses represent a basic step towards a concrete mathematical understanding of new phenomenon that may arise in learning temporal relationships using recurrent architectures.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- On the Generalization Properties of Diffusion ModelsPuheng Li, Zhong Li, Huishuai Zhang, Jiang BianNeurIPS 2023 · 被引用 86 次
- Approximation Rate of the Transformer Architecture for Sequence ModelingHaotian Jiang, Qianxiao LiNeurIPS 2024 · 被引用 32 次
- Understanding the Expressive Power and Mechanisms of Transformer for Sequence ModelingMingze Wang, Weinan ENeurIPS 2024 · 被引用 32 次
- StableSSM: Alleviating the Curse of Memory in State-space Models through Stable ReparameterizationShida Wang, Qianxiao LiICML 2024 · 被引用 25 次
- Approximation Theory of Convolutional Architectures for Time Series ModellingHaotian Jiang, Zhong Li, Qianxiao LiICML 2021 · 被引用 14 次
相关 Paper
- Inverse Approximation Theory for Nonlinear Recurrent Neural NetworksShida Wang, Zhong Li, Qianxiao LiICLR 2024 · 被引用 10 次
- Identifying nonlinear dynamical systems with multiple time scales and long-range dependenciesDominik Schmidt, Georgia Koppe, Zahra Monfared, Max Beutelspacher 等ICLR 2021 · 被引用 41 次
- On the approximation properties of recurrent encoder-decoder architecturesZhong Li, Haotian Jiang, Qianxiao LiICLR 2022 · 被引用 8 次
- Implicit Bias of Linear RNNsMelikasadat Emami, Mojtaba Sahraee-Ardakan, Parthe Pandit, Sundeep Rangan 等ICML 2021 · 被引用 14 次
- Feedback-driven recurrent quantum neural network universalityLukas Gonon, Rodrigo Martínez-Peña, Juan-Pablo OrtegaICLR 2026 · 被引用 8 次
