TidalDecode: Fast and Accurate LLM Decoding with Position Persistent Sparse Attention
Lijie Yang, Zhihao Zhang, Zhuofu Chen, Zikun Li, Zhihao Jia
摘要
Large language models (LLMs) have driven significant advancements across diverse NLP tasks, with long-context models gaining prominence for handling extended inputs. However, the expanding key-value (KV) cache size required by Transformer architectures intensifies the memory constraints, particularly during the decoding phase, creating a significant bottleneck. Existing sparse attention mechanisms designed to address this bottleneck have two limitations: (1) they often fail to reliably identify the most relevant tokens for attention, and (2) they overlook the spatial coherence of token selection across consecutive Transformer layers, which can lead to performance degradation and substantial overhead in token selection. This paper introduces TidalDecode, a simple yet effective algorithm and system for fast and accurate LLM decoding through position persistent sparse attention. TidalDecode leverages the spatial coherence of tokens selected by existing sparse attention methods and introduces a few token selection layers that perform full attention to identify the tokens with the highest attention scores, while all other layers perform sparse attention with the pre-selected tokens. This design enables TidalDecode to substantially reduce the overhead of token selection for sparse attention without sacrificing the quality of the generated results. Evaluation on a diverse set of LLMs and tasks shows that TidalDecode closely matches the generative performance of full attention methods while reducing the LLM decoding latency by up to 2.1x.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Twilight: Adaptive Attention Sparsity with Hierarchical Top- PruningChaofan Lin, Jiaming Tang, Shuo Yang, Hanshuo Wang 等NeurIPS 2025 · 被引用 53 次
- SeerAttention: Self-distilled Attention Gating for Efficient Long-context PrefillingYizhao Gao, Zhichen Zeng, Dayou Du, Shijie Cao 等NeurIPS 2025 · 被引用 12 次
- SlimInfer: Accelerating Long-Context LLM Inference via Dynamic Token PruningLingkun Long, Rubing Yang, Yushi Huang, Desheng Hui 等AAAI 2026 · 被引用 8 次
- LycheeDecode: Accelerating Long-Context LLM Inference via Hybrid-Head Sparse DecodingGang Lin, Dongfang Li, Zhuoen Chen, Yukun Shi 等ICLR 2026 · 被引用 5 次
- QuoKA: Query-Oriented KV Selection for Efficient LLM PrefillDalton Jones, Junyoung Park, Matthew J. Morse, Mingu Lee 等ICLR 2026 · 被引用 4 次
相关 Paper
- Vegas: Self-Speculative Decoding with Verification-Guided Sparse AttentionYikang Yue, Yuqi Xue, Jian HuangICML 2026 · 被引用 2 次
- Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token SelectionDongwon Jo, Beomseok Kang, Jiwon Song, jae-joon kimICML 2026 · 被引用 1 次
- Tactic: Adaptive Sparse Attention with Clustering and Distribution Fitting for Long-Context LLMsKan Zhu, Tian Tang, Qinyu Xu, Zhan Jin 等ICLR 2026 · 被引用 25 次
- TileSparse: Arithmetic-Intensity-Aware Sparse Attention for Compute-Bound LLM DecodingChao Wang, Pengfei Zuo, Zhangyu Chen, Qihui Zhou 等ICML 2026
- Latent-Condensed Transformer for Efficient Long Context ModelingZeng You, Yaofo Chen, Qiuwu Chen, Ying Sun 等ACL 2026
