Long Expressive Memory for Sequence Modeling
T. Konstantin Rusch, Siddhartha Mishra, N. Benjamin Erichson, Michael W. Mahoney
Abstract
We propose a novel method called Long Expressive Memory (LEM) for learning long-term sequential dependencies. LEM is gradient-based, it can efficiently process sequential tasks with very long-term dependencies, and it is sufficiently expressive to be able to learn complicated input-output maps. To derive LEM, we consider a system of multiscale ordinary differential equations, as well as a suitable time-discretization of this system. For LEM, we derive rigorous bounds to show the mitigation of the exploding and vanishing gradients problem, a well-known challenge for gradient-based recurrent sequential learning methods. We also prove that LEM can approximate a large class of dynamical systems to high accuracy. Our empirical results, ranging from image and time-series classification through dynamical systems prediction to speech recognition and language modeling, demonstrate that LEM outperforms state-of-the-art recurrent neural networks, gated recurrent units, and long short-term memory models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext de47443c-8219-4b07-9c0b-e226109a7a47Cited by top-tier papers18
- Generalized Teacher Forcing for Learning Chaotic DynamicsFlorian Hess, Zahra Monfared, Manuel Brenner, Daniel DurstewitzICML 2023 · 67 citations
- Tractable Dendritic RNNs for Reconstructing Nonlinear Dynamical SystemsManuel Brenner, Florian Hess, Jonas M. Mikhaeil, Leonard F. Bereska et al.ICML 2022 · 48 citations
- Improving day-ahead Solar Irradiance Time Series Forecasting by Leveraging Spatio-Temporal ContextOussama Boussif, Ghait Boukachab, Dan Assouline, Stefano Massaroli et al.NeurIPS 2023 · 46 citations
- Inferring stochastic low-rank recurrent neural networks from neural dataMatthijs Pals, A Erdem Sagtekin, Felix Pei, Manuel Glöckler et al.NeurIPS 2024 · 37 citations
- Parallelizing non-linear sequential models over the sequence lengthYi Heng Lim, Qi Zhu, Joshua Selfridge, Muhammad Firmansyah KasimICLR 2024 · 33 citations
Builds on8
- Multiscale Deep Equilibrium ModelsShaojie Bai, Vladlen Koltun, J. Zico KolterNeurIPS 2020 · 272 citations
- Symplectic Recurrent Neural NetworksZhengdao Chen, Jianyu Zhang, Martín Arjovsky, Léon BottouICLR 2020 · 261 citations
- Neural Rough Differential Equations for Long Time SeriesJames Morrill, Cristopher Salvi, Patrick Kidger, James FosterICML 2021 · 176 citations
- Coupled Oscillatory Recurrent Neural Network (coRNN): An accurate and (gradient) stable architecture for learning long time dependenciesT. Konstantin Rusch, Siddhartha MishraICLR 2021 · 121 citations
- Noisy Recurrent Neural NetworksSoon Hoe Lim, N. Benjamin Erichson, Liam Hodgkinson, Michael W. MahoneyNeurIPS 2021 · 77 citations
Related papers
- UnICORNN: A recurrent model for learning very long time dependenciesT. Konstantin Rusch, Siddhartha MishraICML 2021 · 76 citations
- Identifying nonlinear dynamical systems with multiple time scales and long-range dependenciesDominik Schmidt, Georgia Koppe, Zahra Monfared, Max Beutelspacher et al.ICLR 2021 · 41 citations
- RNNs Incrementally Evolving on an Equilibrium Manifold: A Panacea for Vanishing and Exploding Gradients?Anil Kag, Ziming Zhang, Venkatesh SaligramaICLR 2020 · 51 citations
- Recurrent neural networks: vanishing and exploding gradients are not the end of the storyNicolas Zucchet, Antonio OrvietoNeurIPS 2024 · 78 citations
- Explain My Surprise: Learning Efficient Long-Term Memory by predicting uncertain outcomesArtyom Y. Sorokin, Nazar Buzun, Leonid Pugachev, Mikhail BurtsevNeurIPS 2022 · 13 citations
