LAD: Efficient Accelerator for Generative Inference of LLM with Locality Aware Decoding
Haoran Wang, Yuming Li, Haobo Xu, Ying Wang, Liqi Liu, Jun Yang, Yinhe Han
Abstract
Large Language Models (LLMs) have emerged as the cornerstone of content generation applications due to their ability to capture relations between newly generated token and the full preceding context. However, this ability stems from the attention mechanism for decoding that retains the entire generation history as key value cache (KV cache). As the generated sequence lengthens, the KV cache expands, causing a substantial memory access bottleneck. In advanced LLM generation systems running on GPUs, the attention mechanism for decoding accounts for more than 50% of the total inference time when the KV cache length reaches 4096. To address this issue, this paper introduces LAD (Locality Aware Decoding), an LLM generation accelerator with algorithm-hardware enhancements that significantly decrease KV cache access, resulting in considerable speedups and energy savings. A key insight underlying LAD is that when the attention score for a specific position remains fixed over the next several decoding steps, it is unnecessary to repeatedly retrieve the associated key and value at each step to reproduce the computation. Our analysis reveals that numerous positions exhibit notable numerical locality in attention scores through multiple decoding steps. Leveraging these insights, we have designed an innovative attention decoding computation method that decreases the frequency of accessing the key and value for positions demonstrating good locality, all while maintaining decoding accuracy. Extensive experiments show that LAD generates sequences with an average ROUGE-1 similarity of 97% compared to those generated by the original model. When the length of KV cache exceeds 2048, the high configuration of LAD accelerator achieves on average (geomean) speedup and energy efficiency for the attention mechanism compared to the A100 GPU. For end-to-end model inference, it also achieves on average speedup and energy efficiency.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 117db049-0cdb-443e-bbb7-6b8f068fea3dCited by top-tier papers2
- Focus: A Streaming Concentration Architecture for Efficient Vision-Language ModelsChiyue Wei, Cong Guo, Junyao Zhang, Haoxuan Shan et al.HPCA 2026 · 2 citations
- NanoSpec: Accelerating Speculative Decoding using Minimalist In-Context VocabulariesZhiyang Chen, Daliang Xu, Yinyuan Zhang, Chenghua Wang et al.ICML 2026 · 1 citation
Related papers
- LazyAttention: Efficient Retrieval-Augmented Generation with Deferred Positional EncodingHaocheng Xia, Mihir Pamnani, Hanxi Fang, Supawit Chockchowwat et al.ICML 2026
- KVO-LLM: Boosting Long-Context Generation Throughput for Batched LLM InferenceZhenyu Li, Dongxu Lyu, Gang Wang, Yuzhou Chen et al.DAC 2025 · 1 citation
- SnapKV: LLM Knows What You are Looking for Before GenerationYuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh et al.NeurIPS 2024 · 1,019 citations
- InstAttention: In-Storage Attention Offloading for Cost-Effective Long-Context LLM InferenceXiurui Pan, Endian Li, Qiao Li, Shengwen Liang et al.HPCA 2025 · 22 citations
- AdaDecode: Accelerating LLM Decoding with Adaptive Layer ParallelismZhepei Wei, Wei-Lin Chen, Xinyu Zhu, Yu MengICML 2025
