Combining Recurrent, Convolutional, and Continuous-time Models with Linear State Space Layers
Albert Gu, Isys Johnson, Karan Goel, Khaled Saab, Tri Dao, Atri Rudra, Christopher Ré
Abstract
Recurrent neural networks (RNNs), temporal convolutions, and neural differential equations (NDEs) are popular families of deep learning models for time-series data, each with unique strengths and tradeoffs in modeling power and computational efficiency. We introduce a simple sequence model inspired by control systems that generalizes these approaches while addressing their shortcomings. The Linear State-Space Layer (LSSL) maps a sequence by simply simulating a linear continuous-time state-space representation . Theoretically, we show that LSSL models are closely related to the three aforementioned families of models and inherit their strengths. For example, they generalize convolutions to continuous-time, explain common RNN heuristics, and share features of NDEs such as time-scale adaptation. We then incorporate and generalize recent theory on continuous-time memorization to introduce a trainable subset of structured matrices that endow LSSLs with long-range memory. Empirically, stacking LSSL layers into a simple deep neural network obtains state-of-the-art results across time series benchmarks for long dependencies in sequential image classification, real-world healthcare regression tasks, and speech. On a difficult speech classification task with length-16000 sequences, LSSL outperforms prior approaches by 24 accuracy points, and even outperforms baselines that use hand-crafted features on 100x shorter sequences.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 70245875-a5e1-4c5a-a653-e85f3af8df59Cited by top-tier papers221
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu et al.NeurIPS 2024 · 3,199 citations
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang et al.ICML 2024 · 1,725 citations
- Non-stationary Transformers: Exploring the Stationarity in Time Series ForecastingYong Liu, Haixu Wu, Jianmin Wang, Mingsheng LongNeurIPS 2022 · 1,080 citations
- Diagonal State Spaces are as Effective as Structured State SpacesAnkit Gupta, Albert Gu, Jonathan BerantNeurIPS 2022 · 546 citations
Builds on13
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- HiPPO: Recurrent Memory with Optimal Polynomial ProjectionsAlbert Gu, Tri Dao, Stefano Ermon, Atri Rudra et al.NeurIPS 2020 · 1,100 citations
- Neural Controlled Differential Equations for Irregular Time SeriesPatrick Kidger, James Morrill, James Foster, Terry J. LyonsNeurIPS 2020 · 850 citations
- Liquid Time-constant NetworksRamin M. Hasani, Mathias Lechner, Alexander Amini, Daniela Rus et al.AAAI 2021 · 399 citations
- Neural Rough Differential Equations for Long Time SeriesJames Morrill, Cristopher Salvi, Patrick Kidger, James FosterICML 2021 · 176 citations
Related papers
- Liquid Structural State-Space ModelsRamin M. Hasani, Mathias Lechner, Tsun-Hsuan Wang, Makram Chahine et al.ICLR 2023 · 13 citations
- EXIT: Extrapolation and Interpolation-based Neural Controlled Differential Equations for Time-series Classification and ForecastingSheo Yon Jhin, Jaehoon Lee, Minju Jo, Seungji Kook et al.WWW 2022 · 30 citations
- Effectively Modeling Time Series with Simple Discrete State SpacesMichael Zhang, Khaled Kamal Saab, Michael Poli, Tri Dao et al.ICLR 2023 · 14 citations
- State-space models with layer-wise nonlinearity are universal approximators with exponential decaying memoryShida Wang, Beichen XueNeurIPS 2023 · 49 citations
- Deep Latent State Space Models for Time-Series GenerationLinqi Zhou, Michael Poli, Winnie Xu, Stefano Massaroli et al.ICML 2023 · 56 citations
