Parallelizing Legendre Memory Unit Training
Narsimha Reddy Chilkuri, Chris Eliasmith
摘要
Recently, a new recurrent neural network (RNN) named the Legendre Memory Unit (LMU) was proposed and shown to achieve state-of-the-art performance on several benchmark datasets. Here we leverage the linear time-invariant (LTI) memory component of the LMU to construct a simplified variant that can be parallelized during training (and yet executed as an RNN during inference), thus overcoming a well known limitation of training RNNs on GPUs. We show that this reformulation that aids parallelizing, which can be applied generally to any deep network whose recurrent components are linear, makes training up to 200 times faster. Second, to validate its utility, we compare its performance against the original LMU and a variety of published LSTM and transformer networks on seven benchmarks, ranging from psMNIST to sentiment analysis to machine translation. We demonstrate that our models exhibit superior performance on all datasets, often using fewer parameters. For instance, our LMU sets a new state-of-the-art result on psMNIST, and uses half the parameters while outperforming Dis-tilBERT and LSTM models on IMDB sentiment analysis.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
- Combining Recurrent, Convolutional, and Continuous-time Models with Linear State Space LayersAlbert Gu, Isys Johnson, Karan Goel, Khaled Saab 等NeurIPS 2021 · 被引用 1,280 次
- FlexConv: Continuous Kernel Convolutions With Differentiable Kernel SizesDavid W. Romero, Robert-Jan Bruintjes, Jakub Mikolaj Tomczak, Erik J. Bekkers 等ICLR 2022 · 被引用 94 次
- Simplified State Space Layers for Sequence ModelingJimmy T. H. Smith, Andrew Warrington, Scott W. LindermanICLR 2023 · 被引用 78 次
- Deep Latent State Space Models for Time-Series GenerationLinqi Zhou, Michael Poli, Winnie Xu, Stefano Massaroli 等ICML 2023 · 被引用 56 次
它引用的顶会 Paper1
相关 Paper
- LMUFormer: Low Complexity Yet Powerful Spiking Model With Legendre Memory UnitsZeyu Liu, Gourav Datta, Anni Li, Peter Anthony BeerelICLR 2024 · 被引用 18 次
- Parallelizing non-linear sequential models over the sequence lengthYi Heng Lim, Qi Zhu, Joshua Selfridge, Muhammad Firmansyah KasimICLR 2024 · 被引用 33 次
- Linear Dynamical Systems as a Core Computational PrimitiveShiva KaulNeurIPS 2020 · 被引用 8 次
- ParaRNN: Unlocking Parallel Training of Nonlinear RNNs for Large Language ModelsFederico Danieli, Pau Rodríguez, Miguel Sarabia, Xavier Suau 等ICLR 2026 · 被引用 18 次
- Parallel Training of GRU Networks with a Multi-Grid Solver for Long SequencesEuhyun Moon, Eric C. CyrICLR 2022 · 被引用 9 次
