DIAS: Distance-based Attention Sparsity for Ultra-Long-Sequence Transformer with Tree-like Processing-in-Memory Architecture
Zekai Chen, Yiming Chen, Teng Wan, Tianyi Yu, Yu Wang, Huazhong Yang, Xueqing Li
Abstract
Long-context inference has become a central focus in recent self-regressive Transformer research. However, challenges still remain in performing decode stage due to the memory bandwidth bottleneck of attention mechanisms and the substantial memory overhead associated with KV cache. Although attention sparsity has been proposed as a potential solution, conventional sparsity methods that rely on heuristic algorithms often suffer from accuracy degradation when applied to ultra-long sequences. To break through the dilemma between accuracy-performance and bandwidth-capacity, this work proposes DIAS, a distancebased irregular attention sparsity approach with processing-inmemory (PIM) architecture. DIAS employs approximate topK attention (AKAttention) scores through graph-based search to improve inference efficiency while maintaining accuracy. Furthermore, a scalable tree-like PIM (TreePIM) architecture is introduced to achieve both memory capacity and bandwidth improvement by isolating enormous memory access for KV cache into the PIM units. Evaluations on various configurations of DIAS for Longbench with Llama3-405B models with 1 M sequence length show up to 75 times speedup compared with the state-of-the-art LLM accelerator, with accuracy drop of less than . Index Terms-AI and Machine Learning, Architecture & System Design
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 4b78f080-e99d-4507-9453-e26963e8b667Related papers
- SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM PrefillingXiaodong Ji, Hailin Zhang, Fangcheng Fu, Bin CuiICML 2026 · 3 citations
- STARC: Selective Token Access with Remapping and Clustering for Efficient LLM Decoding on PIM SystemsZehao Fan, Yunzhen Liu, Garrett Gagnon, Zhenyu Liu et al.ASPLOS 2026
- Long-Context Modeling with Dynamic Hierarchical Sparse Attention for Memory-Constrained LLM InferenceSiheng Xiong, Joe Zou, Faramarz Fekri, Yae Jee ChoICML 2026
- ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM InferenceHanshi Sun, Li-Wen Chang, Wenlei Bao, Size Zheng et al.ICML 2025
- SparK: Query-Aware Unstructured Sparsity with Recoverable KV Cache Channel PruningHuanxuan Liao, Yixing Xu, Shizhu He, Guanchen Li et al.AAAI 2026 · 3 citations
