Making Every Head Count: Sparse Attention Without the Speed-Performance Trade-off
Mingkuan Zhao, Wentao Hu, Jiayin Wang, Xin Lai, Tianchen Huang, Yuheng Min, Rui Yan, Xiaoyan Zhu
Abstract
The design of Large Language Models (LLMs) has long been hampered by a fundamental conflict within their core attention mechanism: its remarkable expressivity is built upon a computational complexity of O(H·N²) that grows quadratically with the context size (N) and linearly with the number of heads (H). This standard implementation harbors significant computational redundancy, as all heads independently compute attention over the same sequence space. Existing sparse methods, meanwhile, often trade information integrity for computational efficiency. To resolve this efficiency-performance trade-off, we propose SPAttention, whose core contribution is the introduction of a new paradigm we term Principled Structural Sparsity. SPAttention does not merely drop connections but instead reorganizes the computational task by partitioning the total attention workload into balanced, non-overlapping distance bands, assigning each head a unique segment. This approach transforms the multi-head attention mechanism from H independent O(N²) computations into a single, collaborative O(N²) computation, fundamentally reducing complexity by a factor of H. The structured inductive bias compels functional specialization among heads, enabling a more efficient allocation of computational resources from redundant modeling to distinct dependencies across the entire sequence span. Extensive empirical validation on the OLMoE-1B-7B and 0.25B-1.75B model series demonstrates that while delivering an approximately two-fold increase in training throughput, its performance is on par with standard dense attention, even surpassing it on select key metrics, while consistently outperforming representative sparse attention methods including Longformer, Reformer, and BigBird across all evaluation metrics. Our work demonstrates that thoughtfully designed structural sparsity can serve as an effective inductive bias that simultaneously improves both computational efficiency and model performance, opening a new avenue for the architectural design of next-generation, high-performance LLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0dc9c254-e088-4781-8ed1-d28732c38066Cited by top-tier papers4
- Mosaic Pruning: A Hierarchical Framework for Generalizable Pruning of Mixture-of-Experts ModelsWentao Hu, Mingkuan Zhao, Shuangyong Song, Xiaoyan Zhu et al.AAAI 2026 · 4 citations
- Regret Pre-training: Bridging Prior and Posterior Views for Enhanced Knowledge GroundingMingkuan Zhao, Xiayu Sun, Wentao Hu, Suquan Chen et al.ICML 2026
- RaGEP: Rank-aware Geometric Expert Pruning for Mixture-of-Experts Language ModelsWentao Hu, Zeyu Zhu, Mingkuan Zhao, Zhenhua An et al.ICML 2026
- Awakening Dormant Experts: Counterfactual Routing to Mitigate MoE HallucinationsWentao Hu, Yanbo Zhai, Xiaohui Hu, Mingkuan Zhao et al.ACL 2026
Builds on11
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
Related papers
- Low-Rank Approximation for Sparse Attention in Multi-Modal LLMsLin Song, Yukang Chen, Shuai Yang, Xiaohan Ding et al.CVPR 2024 · 8 citations
- Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse AttentionJingyang Yuan, Huazuo Gao, Damai Dai, Junyu Luo et al.ACL 2025 · 334 citations
- FSA: An Alternative Efficient Implementation of Native Sparse Attention KernelRan Yan, Youhe Jiang, Zhuoming Chen, Haohui Mai et al.ICLR 2026 · 10 citations
- RRAtention: Dynamic Block Sparse Attention via Per-Head Round-Robin Shifts for Long-Context InferenceSiran Liu, Guoxia Wang, Sa Wang, Jinle Zeng et al.ACL 2026
- Long-Context Attention Benchmark: From Kernel Efficiency to Distributed Context ParallelismTao Bu, Qiangang Wang, Bowen Zeng, Hanwen Sun et al.ICLR 2026
