Learning and Generalization in RNNs
Abhishek Panigrahi, Navin Goyal
摘要
Simple recurrent neural networks (RNNs) and their more advanced cousins LSTMs etc. have been very successful in sequence modeling. Their theoretical understanding, however, is lacking and has not kept pace with the progress for feedforward networks, where a reasonably complete understanding in the special case of highly overparametrized one-hidden-layer networks has emerged. In this paper, we make progress towards remedying this situation by proving that RNNs can learn functions of sequences. In contrast to the previous work that could only deal with functions of sequences that are sums of functions of individual tokens in the sequence, we allow general functions. Conceptually and technically, we introduce new ideas which enable us to extract information from the hidden state of the RNN in our proofs -- addressing a crucial weakness in previous work. We illustrate our results on some regular language recognition problems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Are Transformers universal approximators of sequence-to-sequence functions?Chulhee Yun, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank J. Reddi 等ICLR 2020 · 被引用 481 次
- Understanding Generalization in Recurrent Neural NetworksZhuozhuo Tu, Fengxiang He, Dacheng TaoICLR 2020 · 被引用 33 次
- On the Ability and Limitations of Transformers to Recognize Formal LanguagesSatwik Bhattamishra, Kabir Ahuja, Navin GoyalEMNLP 2020 · 被引用 7 次
- A Formal Hierarchy of RNN ArchitecturesWilliam Merrill, Gail Weiss, Yoav Goldberg, Roy Schwartz 等ACL 2020 · 被引用 6 次
- The Recurrent Neural Tangent KernelSina Alemohammad, Zichao Wang, Randall Balestriero, Richard G. BaraniukICLR 2021 · 被引用 6 次
相关 Paper
- On the Provable Generalization of Recurrent Neural NetworksLifu Wang, Bo Shen, Bo Hu, Xing CaoNeurIPS 2021 · 被引用 9 次
- What Languages are Easy to Language-Model? A Perspective from Learning Probabilistic Regular LanguagesNadav Borenstein, Anej Svete, Robin Chan, Josef Valvoda 等ACL 2024
- Identifying nonlinear dynamical systems with multiple time scales and long-range dependenciesDominik Schmidt, Georgia Koppe, Zahra Monfared, Max Beutelspacher 等ICLR 2021 · 被引用 41 次
- Learning Useful Representations of Recurrent Neural Network Weight MatricesVincent Herrmann, Francesco Faccio, Jürgen SchmidhuberICML 2024 · 被引用 12 次
- Framing RNN as a kernel method: A neural ODE approachAdeline Fermanian, Pierre Marion, Jean-Philippe Vert, Gérard BiauNeurIPS 2021 · 被引用 34 次
