Making Every Head Count: Sparse Attention Without the Speed-Performance Trade-off
Mingkuan Zhao, Wentao Hu, Jiayin Wang, Xin Lai, Tianchen Huang, Yuheng Min, Rui Yan, Xiaoyan Zhu
摘要
The design of Large Language Models (LLMs) has long been hampered by a fundamental conflict within their core attention mechanism: its remarkable expressivity is built upon a computational complexity of O(H·N²) that grows quadratically with the context size (N) and linearly with the number of heads (H). This standard implementation harbors significant computational redundancy, as all heads independently compute attention over the same sequence space. Existing sparse methods, meanwhile, often trade information integrity for computational efficiency. To resolve this efficiency-performance trade-off, we propose SPAttention, whose core contribution is the introduction of a new paradigm we term Principled Structural Sparsity. SPAttention does not merely drop connections but instead reorganizes the computational task by partitioning the total attention workload into balanced, non-overlapping distance bands, assigning each head a unique segment. This approach transforms the multi-head attention mechanism from H independent O(N²) computations into a single, collaborative O(N²) computation, fundamentally reducing complexity by a factor of H. The structured inductive bias compels functional specialization among heads, enabling a more efficient allocation of computational resources from redundant modeling to distinct dependencies across the entire sequence span. Extensive empirical validation on the OLMoE-1B-7B and 0.25B-1.75B model series demonstrates that while delivering an approximately two-fold increase in training throughput, its performance is on par with standard dense attention, even surpassing it on select key metrics, while consistently outperforming representative sparse attention methods including Longformer, Reformer, and BigBird across all evaluation metrics. Our work demonstrates that thoughtfully designed structural sparsity can serve as an effective inductive bias that simultaneously improves both computational efficiency and model performance, opening a new avenue for the architectural design of next-generation, high-performance LLMs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Mosaic Pruning: A Hierarchical Framework for Generalizable Pruning of Mixture-of-Experts ModelsWentao Hu, Mingkuan Zhao, Shuangyong Song, Xiaoyan Zhu 等AAAI 2026 · 被引用 4 次
- Regret Pre-training: Bridging Prior and Posterior Views for Enhanced Knowledge GroundingMingkuan Zhao, Xiayu Sun, Wentao Hu, Suquan Chen 等ICML 2026
- RaGEP: Rank-aware Geometric Expert Pruning for Mixture-of-Experts Language ModelsWentao Hu, Zeyu Zhu, Mingkuan Zhao, Zhenhua An 等ICML 2026
- Awakening Dormant Experts: Counterfactual Routing to Mitigate MoE HallucinationsWentao Hu, Yanbo Zhai, Xiaohui Hu, Mingkuan Zhao 等ACL 2026
它引用的顶会 Paper11
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
相关 Paper
- Low-Rank Approximation for Sparse Attention in Multi-Modal LLMsLin Song, Yukang Chen, Shuai Yang, Xiaohan Ding 等CVPR 2024 · 被引用 8 次
- Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse AttentionJingyang Yuan, Huazuo Gao, Damai Dai, Junyu Luo 等ACL 2025 · 被引用 334 次
- FSA: An Alternative Efficient Implementation of Native Sparse Attention KernelRan Yan, Youhe Jiang, Zhuoming Chen, Haohui Mai 等ICLR 2026 · 被引用 10 次
- RRAtention: Dynamic Block Sparse Attention via Per-Head Round-Robin Shifts for Long-Context InferenceSiran Liu, Guoxia Wang, Sa Wang, Jinle Zeng 等ACL 2026
- Long-Context Attention Benchmark: From Kernel Efficiency to Distributed Context ParallelismTao Bu, Qiangang Wang, Bowen Zeng, Hanwen Sun 等ICLR 2026
