Flash Inference: Near Linear Time Inference for Long Convolution Sequence Models and Beyond
Costin-Andrei Oncescu, Sanket Purandare, Stratos Idreos, Sham M. Kakade
Abstract
While transformers have been at the core of most recent advancements in sequence generative models, their computational cost remains quadratic in sequence length. Several subquadratic architectures have been proposed to address this computational issue. Some of them, including long convolution sequence models (LCSMs), address this issue at training time but remain quadratic during inference. We propose a method for speeding up LCSMs' exact inference to quasilinear time, identify the key properties that make this possible, and propose a general framework that exploits these. Our approach, inspired by previous work on relaxed polynomial interpolation, is based on a tiling which helps decrease memory movement and share computation. It has the added benefit of allowing for almost complete parallelization across layers of the position-mixing part of the architecture. Empirically, we provide a proof of concept implementation for Hyena, which gets up to 7.8× end-to-end improvement over standard inference by improving up to 110× within the position-mixing part.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext df695328-ad7e-4d5c-87e7-40bcc659790bCited by top-tier papers2
- Log-Linear AttentionHan Guo, Songlin Yang, Tarushii Goel, Eric P. Xing et al.ICLR 2026 · 41 citations
- FutureFill: Fast Generation from Convolutional Sequence ModelsNaman Agarwal, Xinyi Chen, Evan Dogariu, Devan Shah et al.ICLR 2026 · 6 citations
Builds on16
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- Combining Recurrent, Convolutional, and Continuous-time Models with Linear State Space LayersAlbert Gu, Isys Johnson, Karan Goel, Khaled Saab et al.NeurIPS 2021 · 1,280 citations
- On the Parameterization and Initialization of Diagonal State Space ModelsAlbert Gu, Karan Goel, Ankit Gupta, Christopher RéNeurIPS 2022 · 690 citations
- Diagonal State Spaces are as Effective as Structured State SpacesAnkit Gupta, Albert Gu, Jonathan BerantNeurIPS 2022 · 546 citations
- Hyena Hierarchy: Towards Larger Convolutional Language ModelsMichael Poli, Stefano Massaroli, Eric Nguyen, Daniel Y. Fu et al.ICML 2023 · 481 citations
Related papers
- Laughing Hyena Distillery: Extracting Compact Recurrences From ConvolutionsStefano Massaroli, Michael Poli, Daniel Y. Fu, Hermann Kumbong et al.NeurIPS 2023 · 31 citations
- Geometric Hyena Networks for Large-scale Equivariant LearningArtem Moskalev, Mangal Prakash, Junjie Xu, Tianyu Cui et al.ICML 2025
- FlashFFTConv: Efficient Convolutions for Long Sequences with Tensor CoresDaniel Y. Fu, Hermann Kumbong, Eric Nguyen, Christopher RéICLR 2024 · 41 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- Hyena Operator for Fast Sequential RecommendationJiahao Liu, Lin Li, Zhiyuan Li, Kaixi Hu et al.WWW 2026
