Simple Hardware-Efficient Long Convolutions for Sequence Modeling
Daniel Y. Fu, Elliot L. Epstein, Eric Nguyen, Armin W. Thomas, Michael Zhang, Tri Dao, Atri Rudra, Christopher Ré
Abstract
State space models (SSMs) have high performance on long sequence modeling but require sophisticated initialization techniques and specialized implementations for high quality and runtime performance. We study whether a simple alternative can match SSMs in performance and efficiency: directly learning long convolutions over the sequence. We find that a key requirement to achieving high performance is keeping the convolution kernels smooth. We find that simple interventions--such as squashing the kernel weights--result in smooth kernels and recover SSM performance on a range of tasks including the long range arena, image classification, language modeling, and brain data modeling. Next, we develop FlashButterfly, an IO-aware algorithm to improve the runtime performance of long convolutions. FlashButterfly appeals to classic Butterfly decompositions of the convolution to reduce GPU memory IO and increase FLOP utilization. FlashButterfly speeds up convolutions by 2.2, and allows us to train on Path256, a challenging task with sequence length 64K, where we set state-of-the-art by 29.1 points while training 7.2 faster than prior work. Lastly, we introduce an extension to FlashButterfly that learns the coefficients of the Butterfly decomposition, increasing expressivity without increasing runtime. Using this extension, we outperform a Transformer on WikiText103 by 0.2 PPL with 30% fewer parameters.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 67323a45-eb10-4435-897c-7ea8ce4850efCited by top-tier papers28
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space DualityTri Dao, Albert GuICML 2024 · 1,407 citations
- HyenaDNA: Long-Range Genomic Sequence Modeling at Single Nucleotide ResolutionEric Nguyen, Michael Poli, Marjan Faizi, Armin W. Thomas et al.NeurIPS 2023 · 574 citations
- Gated Linear Attention Transformers with Hardware-Efficient TrainingSonglin Yang, Bailin Wang, Yikang Shen, Rameswar Panda et al.ICML 2024 · 390 citations
- Hierarchically Gated Recurrent Neural Network for Sequence ModelingZhen Qin, Songlin Yang, Yiran ZhongNeurIPS 2023 · 152 citations
- Zoology: Measuring and Improving Recall in Efficient Language ModelsSimran Arora, Sabri Eyuboglu, Aman Timalsina, Isys Johnson et al.ICLR 2024 · 140 citations
Builds on27
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series ForecastingHaoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang et al.AAAI 2021 · 7,289 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- HiPPO: Recurrent Memory with Optimal Polynomial ProjectionsAlbert Gu, Tri Dao, Stefano Ermon, Atri Rudra et al.NeurIPS 2020 · 1,100 citations
Related papers
- Hungry Hungry Hippos: Towards Language Modeling with State Space ModelsDaniel Y. Fu, Tri Dao, Khaled Kamal Saab, Armin W. Thomas et al.ICLR 2023 · 117 citations
- VGA: Hardware Accelerator for Scalable Long Sequence Model InferenceSeung Yul Lee, Hyunseung Lee, Jihoon Hong, SangLyul Cho et al.MICRO 2024 · 7 citations
- Never Train from Scratch: Fair Comparison of Long-Sequence Models Requires Data-Driven PriorsIdo Amos, Jonathan Berant, Ankit GuptaICLR 2024 · 39 citations
- Block-State TransformersJonathan Pilault, Mahan Fathi, Orhan Firat, Chris Pal et al.NeurIPS 2023 · 33 citations
- What Makes Convolutional Models Great on Long Sequence Modeling?Yuhong Li, Tianle Cai, Yi Zhang, Deming Chen et al.ICLR 2023 · 20 citations
