Scaling Attention via Feature Sparsity
Yan Xie, Tiansheng Wen, Tangda Huang, Bo Chen, Chenyu You, Stefanie Jegelka, Yifei Wang
Abstract
Scaling Transformers to ultra-long contexts is bottlenecked by the O(n 2 d) cost of self-attention. Existing methods reduce this cost along the sequence axis through local windows, kernel approximations, or token-level sparsity, but these approaches consistently degrade accuracy. In this paper, we instead explore an orthogonal axis: feature sparsity. We propose Sparse Feature Attention (SFA), where queries and keys are represented as k-sparse codes that preserve highdimensional expressivity while reducing the cost of attention from Θ(n 2 d) to Θ(n 2 k 2 /d). To make this efficient at scale, we introduce FlashSFA, an IOaware kernel that extends FlashAttention to operate directly on sparse overlaps without materializing dense score matrices. Across GPT-2 and Qwen3 pretraining, SFA matches dense baselines while improving speed by up to 2.5× and reducing FLOPs and KV-cache by nearly 50%. On synthetic and downstream benchmarks, SFA preserves retrieval accuracy and robustness at long contexts, outperforming short-embedding baselines that collapse feature diversity. These results establish feature-level sparsity as a complementary and underexplored axis for efficient attention, enabling Transformers to scale to ordersof-magnitude longer contexts with minimal quality loss. Code is available at https://github.com/YannX1e/Sparse-Feature-Attention . (a) Latency comparison (b) FLOPs & KV-cache comparison * Equal Contribution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 64138a69-4ad2-444f-8fce-c7c1a4316844Cited by top-tier papers2
- No More K-means: Single-Stage Sparse Coding for Efficient Multi-Vector RetrievalLixuan Guo, Yifei Wang, Tiansheng Wen, Aosong Feng et al.ICML 2026
- Supervise Less, See More: Training-free Nuclear Instance Segmentation with Prototype-Guided PromptingWen Zhang, Qin Ren, Wenjing Liu, Haibin Ling et al.ICML 2026
Related papers
- Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token SelectionDongwon Jo, Beomseok Kang, Jiwon Song, jae-joon kimICML 2026 · 1 citation
- FSA: An Alternative Efficient Implementation of Native Sparse Attention KernelRan Yan, Youhe Jiang, Zhuoming Chen, Haohui Mai et al.ICLR 2026 · 10 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Fast Attention Over Long Sequences With Dynamic Sparse Flash AttentionMatteo Pagliardini, Daniele Paliotta, Martin Jaggi, François FleuretNeurIPS 2023 · 26 citations
- FlashMask: Efficient and Rich Mask Extension of FlashAttentionGuoxia Wang, Jinle Zeng, Xiyuan Xiao, Siming Wu et al.ICLR 2025
