Memory-Efficient KV Cache Optimization for Large Language Model Inference at the Edge
Chi Zhang, Haisheng Tan, Haotian Pan, Yang Xu, Haohua Du, Li Zhang, Xiaoming Fu
Abstract
The deployment of large language models (LLMs) on edge devices offers notable advantages in privacy preservation, low-latency response, and operational autonomy, but faces significant challenges due to the memory-intensive key-value (KV) cache required for autoregressive decoding. Sparse attention methods alleviate computational costs but typically fail to preserve the generation quality and inference efficiency under memory constraints, which becomes particularly acute on edge platforms. In this paper, we first formalize and prove the NP-hardness of the KV cache replacement problem under top-k sparse attention, where k is the number of pages involved in each attention calculation. Then, we propose LiteKV, a lightweight memory-efficient online KV cache replacement algorithm tailored for edge-based LLM inference. LiteKV introduces a (1 + ϵ)-relaxation strategy that tolerates minimal precision loss in attention scores, allowing significant improvements in cache utilization. Moreover, LiteKV is orthogonal to existing sparse attention and quantization methods, enabling scalable, high-performance LLM inference under stringent memory budgets. We provide theoretical guarantees on LiteKV’s competitive ratio under realistic attention score distributions. Experiments on real-world workloads demonstrate that LiteKV reduces memory usage by up to 75% compared to state-of-the-art memory-unconstrained baselines, while preserving comparable generation quality and 95% of their decoding efficiency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on18
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- SnapKV: LLM Knows What You are Looking for Before GenerationYuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh et al.NeurIPS 2024 · 1,019 citations
- H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language ModelsZhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen et al.NeurIPS 2023 · 1,003 citations
- Compressive Transformers for Long-Range Sequence ModellingJack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Chloe Hillier et al.ICLR 2020 · 833 citations
- KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache QuantizationColeman Hooper, Sehoon Kim, Hiva Mohammadzadeh, Michael W. Mahoney et al.NeurIPS 2024 · 738 citations
Related papers
- Kelle: Co-design KV Caching and eDRAM for Efficient LLM Serving in Edge ComputingTianhua Xia, Sai Qian ZhangMICRO 2025 · 2 citations
- Efficient Low Rank Attention for Long-Context Inference in Large Language ModelsTenghui Li, Guoxu Zhou, Xuyang Zhao, Yuning Qiu et al.NeurIPS 2025 · 4 citations
- KeepKV: Achieving Periodic Lossless KV Cache Compression for Efficient LLM InferenceYuxuan Tian, Zihan Wang, Yebo Peng, Aomufei Yuan et al.AAAI 2026
- Value-Guided KV Compression for LLMs via Approximated CUR DecompositionAyan Sengupta, Siddhant Chaudhary, Tanmoy ChakrabortyNeurIPS 2025 · 6 citations
- Get More with LESS: Synthesizing Recurrence with KV Cache Compression for Efficient LLM InferenceHarry Dong, Xinyu Yang, Zhenyu Zhang, Zhangyang Wang et al.ICML 2024 · 84 citations
