AnchorAttention: Difference-Aware Sparse Attention with Stripe Granularity
Yu Zhang, Dong Guo, Fang Wu, Guoliang Zhu, Dian Ding, Yiming Zhang
Abstract
Large Language Models (LLMs) with extended context lengths face significant computational challenges during the pre-filling phase, primarily due to the quadratic complexity of selfattention. Existing methods typically employ dynamic pattern matching and block-sparse low-level implementations. However, their reliance on local information for pattern identification fails to capture global contexts, and the coarse granularity of blocks leads to persistent internal sparsity, resulting in subop-
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cf07947c-5ab3-4874-9a10-466e26eaa5c5Cited by top-tier papers3
- SparseD: Sparse Attention for Diffusion Language ModelsZeqing Wang, Gongfan Fang, Xinyin Ma, Xingyi Yang et al.ICLR 2026 · 18 citations
- VecAttention: Vector-wise Sparse Attention for Accelerating Long Context InferenceAnmin Liu, Ruixuan Yang, Huiqiang Jiang, Bin Lin et al.CVPR 2026 · 4 citations
- S2O: Early Stopping for Sparse Attention via Online PermutationYu Zhang, Songwei Liu, Chenqian Yan, Sheng Lin et al.ACL 2026
Builds on13
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- SnapKV: LLM Knows What You are Looking for Before GenerationYuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh et al.NeurIPS 2024 · 1,019 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language ModelsZhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen et al.NeurIPS 2023 · 1,003 citations
Related papers
- Sparser Block-Sparse Attention via Token PermutationXinghao Wang, Pengyu Wang, Dong Zhang, Chenkun Tan et al.ICML 2026 · 2 citations
- FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence InferenceXunhao Lai, Jianqiao Lu, Yao Luo, Yiyuan Ma et al.ICLR 2025
- MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse AttentionHuiqiang Jiang, Yucheng Li, Chengruidong Zhang, Qianhui Wu et al.NeurIPS 2024 · 479 citations
- SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM PrefillingXiaodong Ji, Hailin Zhang, Fangcheng Fu, Bin CuiICML 2026 · 3 citations
- UltraAttn: Efficiently Parallelizing Attention through Hierarchical Context-TilingHaoyu Yang, Zan Zong, Yuyang Jin, Kinman Lei et al.SC 2025 · 1 citation
