FutureFill: Fast Generation from Convolutional Sequence Models
Naman Agarwal, Xinyi Chen, Evan Dogariu, Devan Shah, Hubert Strauss, Vladimir Feinberg, Daniel Suo, Peter L. Bartlett, Elad Hazan
摘要
We address the challenge of efficient auto-regressive generation in sequence prediction models by introducing FutureFill-a general-purpose fast generation method for any sequence prediction algorithm based on convolutional operators. Future-Fill reduces generation time from quadratic to quasilinear in the context length. Moreover, when generating from a prompt, it requires a prefill cache whose size grows only with the number of tokens to be generated-often much smaller than the caches required by standard convolutional or attention-based models. We validate our theoretical claims with experiments on synthetic tasks and demonstrate substantial efficiency gains when generating from a deep convolutional sequence prediction model.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Universal Sequence PreconditioningAnnie Marsden, Elad HazanNeurIPS 2025 · 被引用 6 次
- Universal Learning of Nonlinear DynamicsEvan Dogariu, Anand Brahmbhatt, Elad HazanICML 2026 · 被引用 5 次
- A New Approach to Controlling Linear Dynamical SystemsAnand Paresh Brahmbhatt, Gon Buzaglo, Sofiia Druchyna, Elad HazanICLR 2026 · 被引用 4 次
- SpectraLDS: Provable Distillation for Linear Dynamical SystemsDevan Shah, Shlomo Fortgang, Sofiia Druchyna, Elad HazanNeurIPS 2025 · 被引用 1 次
- Efficient Spectral Control of Partially Observed Linear Dynamical SystemsAnand Brahmbhatt, Gon Buzaglo, Sofiia Druchyna, Elad HazanNeurIPS 2025
它引用的顶会 Paper18
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- Combining Recurrent, Convolutional, and Continuous-time Models with Linear State Space LayersAlbert Gu, Isys Johnson, Karan Goel, Khaled Saab 等NeurIPS 2021 · 被引用 1,280 次
- Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationOfir Press, Noah A. Smith, Mike LewisICLR 2022 · 被引用 1,168 次
相关 Paper
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 被引用 2,665 次
- RefreshKV: Updating Small KV Cache During Long-form GenerationFangyuan Xu, Tanya Goyal, Eunsol ChoiACL 2025 · 被引用 6 次
- FastVAR: Linear Visual Autoregressive Modeling Via Cached Token PruningHang Guo, Yawei Li, Taolin Zhang, Jiangshan Wang 等ICCV 2025 · 被引用 5 次
- PKAS: Predictive KVCache-Aware Scheduling for Faster LLM and Transformer InferencesJie Ye, Avinash Maurya, Krishna Teja Chitty-Venkata, Bogdan Nicolae 等HPDC 2026
- Autoregression with Self-Token PredictionDengsheng Chen, Yangming Shi, Enhua WuICML 2026
