Long Expressive Memory for Sequence Modeling
T. Konstantin Rusch, Siddhartha Mishra, N. Benjamin Erichson, Michael W. Mahoney
摘要
We propose a novel method called Long Expressive Memory (LEM) for learning long-term sequential dependencies. LEM is gradient-based, it can efficiently process sequential tasks with very long-term dependencies, and it is sufficiently expressive to be able to learn complicated input-output maps. To derive LEM, we consider a system of multiscale ordinary differential equations, as well as a suitable time-discretization of this system. For LEM, we derive rigorous bounds to show the mitigation of the exploding and vanishing gradients problem, a well-known challenge for gradient-based recurrent sequential learning methods. We also prove that LEM can approximate a large class of dynamical systems to high accuracy. Our empirical results, ranging from image and time-series classification through dynamical systems prediction to speech recognition and language modeling, demonstrate that LEM outperforms state-of-the-art recurrent neural networks, gated recurrent units, and long short-term memory models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- Generalized Teacher Forcing for Learning Chaotic DynamicsFlorian Hess, Zahra Monfared, Manuel Brenner, Daniel DurstewitzICML 2023 · 被引用 67 次
- Tractable Dendritic RNNs for Reconstructing Nonlinear Dynamical SystemsManuel Brenner, Florian Hess, Jonas M. Mikhaeil, Leonard F. Bereska 等ICML 2022 · 被引用 48 次
- Improving day-ahead Solar Irradiance Time Series Forecasting by Leveraging Spatio-Temporal ContextOussama Boussif, Ghait Boukachab, Dan Assouline, Stefano Massaroli 等NeurIPS 2023 · 被引用 46 次
- Inferring stochastic low-rank recurrent neural networks from neural dataMatthijs Pals, A Erdem Sagtekin, Felix Pei, Manuel Glöckler 等NeurIPS 2024 · 被引用 37 次
- Parallelizing non-linear sequential models over the sequence lengthYi Heng Lim, Qi Zhu, Joshua Selfridge, Muhammad Firmansyah KasimICLR 2024 · 被引用 33 次
它引用的顶会 Paper8
- Multiscale Deep Equilibrium ModelsShaojie Bai, Vladlen Koltun, J. Zico KolterNeurIPS 2020 · 被引用 272 次
- Symplectic Recurrent Neural NetworksZhengdao Chen, Jianyu Zhang, Martín Arjovsky, Léon BottouICLR 2020 · 被引用 261 次
- Neural Rough Differential Equations for Long Time SeriesJames Morrill, Cristopher Salvi, Patrick Kidger, James FosterICML 2021 · 被引用 176 次
- Coupled Oscillatory Recurrent Neural Network (coRNN): An accurate and (gradient) stable architecture for learning long time dependenciesT. Konstantin Rusch, Siddhartha MishraICLR 2021 · 被引用 121 次
- Noisy Recurrent Neural NetworksSoon Hoe Lim, N. Benjamin Erichson, Liam Hodgkinson, Michael W. MahoneyNeurIPS 2021 · 被引用 77 次
相关 Paper
- UnICORNN: A recurrent model for learning very long time dependenciesT. Konstantin Rusch, Siddhartha MishraICML 2021 · 被引用 76 次
- Identifying nonlinear dynamical systems with multiple time scales and long-range dependenciesDominik Schmidt, Georgia Koppe, Zahra Monfared, Max Beutelspacher 等ICLR 2021 · 被引用 41 次
- RNNs Incrementally Evolving on an Equilibrium Manifold: A Panacea for Vanishing and Exploding Gradients?Anil Kag, Ziming Zhang, Venkatesh SaligramaICLR 2020 · 被引用 51 次
- Recurrent neural networks: vanishing and exploding gradients are not the end of the storyNicolas Zucchet, Antonio OrvietoNeurIPS 2024 · 被引用 78 次
- Explain My Surprise: Learning Efficient Long-Term Memory by predicting uncertain outcomesArtyom Y. Sorokin, Nazar Buzun, Leonid Pugachev, Mikhail BurtsevNeurIPS 2022 · 被引用 13 次
