Parallelizing non-linear sequential models over the sequence length
Yi Heng Lim, Qi Zhu, Joshua Selfridge, Muhammad Firmansyah Kasim
摘要
Sequential models, such as Recurrent Neural Networks and Neural Ordinary Differential Equations, have long suffered from slow training due to their inherent sequential nature. For many years this bottleneck has persisted, as many thought sequential models could not be parallelized. We challenge this long-held belief with our parallel algorithm that accelerates GPU evaluation of sequential models by up to 3 orders of magnitude faster without compromising output accuracy. The algorithm does not need any special structure in the sequential models' architecture, making it applicable to a wide range of architectures. Using our method, training sequential models can be more than 10 times faster than the common sequential method without any meaningful difference in the training results. Leveraging this accelerated training, we discovered the efficacy of the Gated Recurrent Unit in a long time series classification problem with 17k time samples. By overcoming the training bottleneck, our work serves as the first step to unlock the potential of non-linear sequential models for long sequence problems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- It's All Connected: A Journey Through Test-Time Memorization, Attentional Bias, Retention, and Online OptimizationAli Behrouz, Meisam Razaviyayn, Peilin Zhong, Vahab MirrokniICLR 2026 · 被引用 63 次
- ATLAS: Learning to Optimally Memorize the Context at Test TimeAli Behrouz, Zeman Li, Praneeth Kacham, Majid Daliri 等ICML 2026 · 被引用 57 次
- Towards Scalable and Stable Parallelization of Nonlinear RNNsXavier Gonzalez, Andrew Warrington, Jimmy T. H. Smith, Scott W. LindermanNeurIPS 2024 · 被引用 47 次
- Accelerating Parallel Sampling of Diffusion ModelsZhiwei Tang, Jiasheng Tang, Hao Luo, Fan Wang 等ICML 2024 · 被引用 30 次
- Structured Linear CDEs: Maximally Expressive and Parallel-in-Time Sequence ModelsBenjamin Walker, Lingyi Yang, Nicola Muca Cirone, Cristopher Salvi 等NeurIPS 2025 · 被引用 21 次
它引用的顶会 Paper22
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
- Neural Controlled Differential Equations for Irregular Time SeriesPatrick Kidger, James Morrill, James Foster, Terry J. LyonsNeurIPS 2020 · 被引用 850 次
- On the Parameterization and Initialization of Diagonal State Space ModelsAlbert Gu, Karan Goel, Ankit Gupta, Christopher RéNeurIPS 2022 · 被引用 690 次
- Resurrecting Recurrent Neural Networks for Long SequencesAntonio Orvieto, Samuel L. Smith, Albert Gu, Anushan Fernando 等ICML 2023 · 被引用 474 次
- Multiscale Deep Equilibrium ModelsShaojie Bai, Vladlen Koltun, J. Zico KolterNeurIPS 2020 · 被引用 272 次
相关 Paper
- Parallel Training of GRU Networks with a Multi-Grid Solver for Long SequencesEuhyun Moon, Eric C. CyrICLR 2022 · 被引用 9 次
- Parallelizing Legendre Memory Unit TrainingNarsimha Reddy Chilkuri, Chris EliasmithICML 2021 · 被引用 47 次
- CKConv: Continuous Kernel Convolution For Sequential DataDavid W. Romero, Anna Kuzina, Erik J. Bekkers, Jakub Mikolaj Tomczak 等ICLR 2022 · 被引用 149 次
- Fast Saturating Gate for Learning Long Time Scales with Recurrent Neural NetworksKentaro Ohno, Sekitoshi Kanai, Yasutoshi IdaAAAI 2023 · 被引用 1 次
- DeepPCR: Parallelizing Sequential Operations in Neural NetworksFederico Danieli, Miguel Sarabia, Xavier Suau Cuadros, Pau Rodríguez 等NeurIPS 2023 · 被引用 13 次
