Sparser Block-Sparse Attention via Token Permutation
Xinghao Wang, Pengyu Wang, Dong Zhang, Chenkun Tan, Shaojun Zhou, Zhaoxiang Liu, Shiguo Lian, Fangxu Liu, Kai Song, Xipeng Qiu
Abstract
Scaling the context length of large language models (LLMs) offers significant benefits but is computationally expensive. This expense stems primarily from the self-attention mechanism, whose complexity with respect to sequence length presents a major bottleneck for both memory and latency. Fortunately, the attention matrix is often sparse, particularly for long sequences, suggesting an opportunity for optimization. Block-sparse attention has emerged as a promising solution that partitions sequences into blocks and skips computation for a subset of these blocks. However, the effectiveness of this method is highly dependent on the underlying attention patterns, which can lead to sub-optimal block-level sparsity. For instance, important key tokens for queries within a single block may be scattered across numerous other blocks, leading to computational redundancy. In this work, we propose Permuted Block-Sparse Attention (PBS-Attn), a plug-and-play method that leverages the permutation properties of attention to increase block-level sparsity and enhance the computational efficiency of LLM prefilling. We conduct comprehensive experiments on challenging long-context datasets, demonstrating that PBS-Attn consistently outperforms existing block-sparse attention methods in model accuracy and closely matches the full attention baseline. Powered by our custom permuted-FlashAttention kernels, PBS-Attn achieves an end-to-end speedup of up to in long-context prefilling, confirming its practical viability. Code available at https://github.com/xinghaow99/pbs-attn.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2bce517a-169c-4a03-a338-ca160cb6bd95Cited by top-tier papers2
- Prism: Spectral-Aware Block-Sparse AttentionXinghao Wang, Pengyu Wang, Xiaoran Liu, Fangxu Liu et al.ICML 2026 · 3 citations
- S2O: Early Stopping for Sparse Attention via Online PermutationYu Zhang, Songwei Liu, Chenqian Yan, Sheng Lin et al.ACL 2026
Builds on19
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationOfir Press, Noah A. Smith, Mike LewisICLR 2022 · 1,168 citations
Related papers
- SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM PrefillingXiaodong Ji, Hailin Zhang, Fangcheng Fu, Bin CuiICML 2026 · 3 citations
- ProxyAttn: Guided Sparse Attention via Representative HeadsYixuan Wang, Huang He, Siqi Bao, Hua Wu et al.ICLR 2026 · 9 citations
- SeerAttention: Self-distilled Attention Gating for Efficient Long-context PrefillingYizhao Gao, Zhichen Zeng, Dayou Du, Shijie Cao et al.NeurIPS 2025 · 12 citations
- APB: Accelerating Distributed Long-Context Inference by Passing Compressed Context Blocks across GPUsYuxiang Huang, Mingye Li, Xu Han, Chaojun Xiao et al.ACL 2025
- Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token SelectionDongwon Jo, Beomseok Kang, Jiwon Song, jae-joon kimICML 2026 · 1 citation
