Parallelizing Legendre Memory Unit Training
Narsimha Reddy Chilkuri, Chris Eliasmith
Abstract
Recently, a new recurrent neural network (RNN) named the Legendre Memory Unit (LMU) was proposed and shown to achieve state-of-the-art performance on several benchmark datasets. Here we leverage the linear time-invariant (LTI) memory component of the LMU to construct a simplified variant that can be parallelized during training (and yet executed as an RNN during inference), thus overcoming a well known limitation of training RNNs on GPUs. We show that this reformulation that aids parallelizing, which can be applied generally to any deep network whose recurrent components are linear, makes training up to 200 times faster. Second, to validate its utility, we compare its performance against the original LMU and a variety of published LSTM and transformer networks on seven benchmarks, ranging from psMNIST to sentiment analysis to machine translation. We demonstrate that our models exhibit superior performance on all datasets, often using fewer parameters. For instance, our LMU sets a new state-of-the-art result on psMNIST, and uses half the parameters while outperforming Dis-tilBERT and LSTM models on IMDB sentiment analysis.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4d17afa5-2b5c-4774-b39c-757547ea686aCited by top-tier papers6
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- Combining Recurrent, Convolutional, and Continuous-time Models with Linear State Space LayersAlbert Gu, Isys Johnson, Karan Goel, Khaled Saab et al.NeurIPS 2021 · 1,280 citations
- FlexConv: Continuous Kernel Convolutions With Differentiable Kernel SizesDavid W. Romero, Robert-Jan Bruintjes, Jakub Mikolaj Tomczak, Erik J. Bekkers et al.ICLR 2022 · 94 citations
- Simplified State Space Layers for Sequence ModelingJimmy T. H. Smith, Andrew Warrington, Scott W. LindermanICLR 2023 · 78 citations
- Deep Latent State Space Models for Time-Series GenerationLinqi Zhou, Michael Poli, Winnie Xu, Stefano Massaroli et al.ICML 2023 · 56 citations
Builds on1
Related papers
- LMUFormer: Low Complexity Yet Powerful Spiking Model With Legendre Memory UnitsZeyu Liu, Gourav Datta, Anni Li, Peter Anthony BeerelICLR 2024 · 18 citations
- Parallelizing non-linear sequential models over the sequence lengthYi Heng Lim, Qi Zhu, Joshua Selfridge, Muhammad Firmansyah KasimICLR 2024 · 33 citations
- Linear Dynamical Systems as a Core Computational PrimitiveShiva KaulNeurIPS 2020 · 8 citations
- ParaRNN: Unlocking Parallel Training of Nonlinear RNNs for Large Language ModelsFederico Danieli, Pau Rodríguez, Miguel Sarabia, Xavier Suau et al.ICLR 2026 · 18 citations
- Parallel Training of GRU Networks with a Multi-Grid Solver for Long SequencesEuhyun Moon, Eric C. CyrICLR 2022 · 9 citations
