Long-range Sequence Modeling with Predictable Sparse Attention
Yimeng Zhuang, Jing Zhang, Mei Tu
Abstract
Self-attention mechanism has been shown to be an effective approach for capturing global context dependencies in sequence modeling, but it suffers from quadratic complexity in time and memory usage. Due to the sparsity of the attention matrix, much computation is redundant. Therefore, in this paper, we design an efficient Transformer architecture, named Fourier Sparse Attention for Transformer (FSAT), for fast long-range sequence modeling. We provide a brand-new perspective for constructing sparse attention matrix, i.e. making the sparse attention matrix predictable. Two core sub-modules are: (1) A fast Fourier transform based hidden state cross module, which captures and pools L^2 semantic combinations in O(LL) time complexity. (2) A sparse attention matrix estimation module, which predicts dominant elements of an attention matrix based on the output of the previous hidden state cross module. By reparameterization and gradient truncation, FSAT successfully learned the index of dominant elements. The overall complexity about the sequence length is reduced from O(L^2) to O(LL). Extensive experiments (natural language, vision, and math) show that FSAT remarkably outperforms the standard multi-head attention and its variants in various long-sequence tasks with low computational costs, and achieves new state-of-the-art results on the Long Range Arena benchmark.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4a4cdc69-a7f2-4889-951c-8c55ecf949aeCited by top-tier papers3
- FSUIE: A Novel Fuzzy Span Mechanism for Universal Information ExtractionTianshuo Peng, Zuchao Li, Lefei Zhang, Bo Du et al.ACL 2023 · 6 citations
- Improving the Robustness of Transformer-based Large Language Models with Dynamic AttentionLujia Shen, Yuwen Pu, Shouling Ji, Changjiang Li et al.NDSS 2024
- Caracal: Causal Architecture via Spectral MixingBINGZHENG GAN, Tianyi Zhang, LI YUSU, Jing Huang et al.ICML 2026
Builds on12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
Related papers
- Diffuser: Efficient Transformers with Multi-Hop Attention Diffusion for Long SequencesAosong Feng, Irene Li, Yuang Jiang, Rex YingAAAI 2023 · 20 citations
- H-Transformer-1D: Fast One-Dimensional Hierarchical Attention for SequencesZhenhai Zhu, Radu SoricutACL 2021
- Composite Slice Transformer: An Efficient Transformer with Composition of Multi-Scale Multi-Range AttentionsMingu Lee, Saurabh Pitre, Tianyu Jiang, Pierre-David Letourneau et al.ICLR 2023
- SWAT: Scalable and Efficient Window Attention-based Transformers Acceleration on FPGAsZhenyu Bai, Pranav Dangi, Huize Li, Tulika MitraDAC 2024 · 12 citations
- Long-Short Transformer: Efficient Transformers for Language and VisionChen Zhu, Wei Ping, Chaowei Xiao, Mohammad Shoeybi et al.NeurIPS 2021 · 180 citations
