From Logical to Computational Sparsity: Structure-Aware Block-Sparse Attention for Long-Code Completion
Yanli Wang, Yanlin Wang, Bowen Zhang, Yiwei Zhang, Daya Guo, Jiachi Chen, Hongyu Zhang, Zibin Zheng
Abstract
Code Large Language Models face critical Time-To-First-Token (TTFT) latency challenges when handling long code completion due to the quadratic complexity (O(n 2 )) of attention mechanisms. While existing sparse attention methods attempt to address this issue, they suffer from three key limitations: (1) general sparse patterns cause excessive accuracy degradation without considering code structure, (2) code-specific methods achieve only logical sparsity without actual computational speedup, and (3) limited adaptation to complex scenarios such as repository-level completion. We propose SabreCoder, a training-free Structure-aware block-sparse attention mechanism that bridges the gap between logical and computational sparsity. SabreCoder parses code into semantic chunks, constructs chunklevel sparse patterns through dependency analysis and similarity matching, and maps them to GPU-friendly block-sparse formats. Extensive experiments on LCC and CrossCodeEval benchmarks demonstrate that SabreCoder reduces TTFT by 45-55% while maintaining accuracy within 3% of dense attention. Global Attn Intra-Chunk Attn Dependency Attn Similarity Attn Block-Level Attn Code Chunks GPU Blocks Mapping Code File Code Chunks Dependency Graph Code Analysis 2 X 2 Kernel Size Apply Sparse Chunk-Level Sparse Block-Level Sparse Mapping from flextls.protocol import Protocol from flextls.field import VectorUInt16Field from flextls.field import ServerNameListField class Extension(Protocol): def init(self, **kwargs): Protocol.init(self, **kwargs) ...
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f72ebcb0-b0f6-49f4-9cf0-e89cccd01adbBuilds on23
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
Related papers
- SAS: Sparse Attention Synthesizer for Efficient Language Model InferenceYuan Zhou, Shaojie Xiang, Lingfan Yu, Zhenyu Song et al.EuroSys 2026
- Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token SelectionDongwon Jo, Beomseok Kang, Jiwon Song, jae-joon kimICML 2026 · 1 citation
- LongCoder: A Long-Range Pre-trained Language Model for Code CompletionDaya Guo, Canwen Xu, Nan Duan, Jian Yin et al.ICML 2023 · 150 citations
- SparseD: Sparse Attention for Diffusion Language ModelsZeqing Wang, Gongfan Fang, Xinyin Ma, Xingyi Yang et al.ICLR 2026 · 18 citations
- A Unified Sparse Attention via Multi-Granularity CompressionSiran Liu, Zheng Cao, Yongchao HeICML 2026
