FutureFill: Fast Generation from Convolutional Sequence Models
Naman Agarwal, Xinyi Chen, Evan Dogariu, Devan Shah, Hubert Strauss, Vladimir Feinberg, Daniel Suo, Peter L. Bartlett, Elad Hazan
Abstract
We address the challenge of efficient auto-regressive generation in sequence prediction models by introducing FutureFill-a general-purpose fast generation method for any sequence prediction algorithm based on convolutional operators. Future-Fill reduces generation time from quadratic to quasilinear in the context length. Moreover, when generating from a prompt, it requires a prefill cache whose size grows only with the number of tokens to be generated-often much smaller than the caches required by standard convolutional or attention-based models. We validate our theoretical claims with experiments on synthetic tasks and demonstrate substantial efficiency gains when generating from a deep convolutional sequence prediction model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Universal Sequence PreconditioningAnnie Marsden, Elad HazanNeurIPS 2025 · 6 citations
- Universal Learning of Nonlinear DynamicsEvan Dogariu, Anand Brahmbhatt, Elad HazanICML 2026 · 5 citations
- A New Approach to Controlling Linear Dynamical SystemsAnand Paresh Brahmbhatt, Gon Buzaglo, Sofiia Druchyna, Elad HazanICLR 2026 · 4 citations
- SpectraLDS: Provable Distillation for Linear Dynamical SystemsDevan Shah, Shlomo Fortgang, Sofiia Druchyna, Elad HazanNeurIPS 2025 · 1 citation
- Efficient Spectral Control of Partially Observed Linear Dynamical SystemsAnand Brahmbhatt, Gon Buzaglo, Sofiia Druchyna, Elad HazanNeurIPS 2025
Builds on18
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Combining Recurrent, Convolutional, and Continuous-time Models with Linear State Space LayersAlbert Gu, Isys Johnson, Karan Goel, Khaled Saab et al.NeurIPS 2021 · 1,280 citations
- Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationOfir Press, Noah A. Smith, Mike LewisICLR 2022 · 1,168 citations
Related papers
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- RefreshKV: Updating Small KV Cache During Long-form GenerationFangyuan Xu, Tanya Goyal, Eunsol ChoiACL 2025 · 6 citations
- FastVAR: Linear Visual Autoregressive Modeling Via Cached Token PruningHang Guo, Yawei Li, Taolin Zhang, Jiangshan Wang et al.ICCV 2025 · 5 citations
- PKAS: Predictive KVCache-Aware Scheduling for Faster LLM and Transformer InferencesJie Ye, Avinash Maurya, Krishna Teja Chitty-Venkata, Bogdan Nicolae et al.HPDC 2026
- Autoregression with Self-Token PredictionDengsheng Chen, Yangming Shi, Enhua WuICML 2026
