Memory-Efficient KV Cache Optimization for Large Language Model Inference at the Edge
Chi Zhang, Haisheng Tan, Haotian Pan, Yang Xu, Haohua Du, Li Zhang, Xiaoming Fu
摘要
The deployment of large language models (LLMs) on edge devices offers notable advantages in privacy preservation, low-latency response, and operational autonomy, but faces significant challenges due to the memory-intensive key-value (KV) cache required for autoregressive decoding. Sparse attention methods alleviate computational costs but typically fail to preserve the generation quality and inference efficiency under memory constraints, which becomes particularly acute on edge platforms. In this paper, we first formalize and prove the NP-hardness of the KV cache replacement problem under top-k sparse attention, where k is the number of pages involved in each attention calculation. Then, we propose LiteKV, a lightweight memory-efficient online KV cache replacement algorithm tailored for edge-based LLM inference. LiteKV introduces a (1 + ϵ)-relaxation strategy that tolerates minimal precision loss in attention scores, allowing significant improvements in cache utilization. Moreover, LiteKV is orthogonal to existing sparse attention and quantization methods, enabling scalable, high-performance LLM inference under stringent memory budgets. We provide theoretical guarantees on LiteKV’s competitive ratio under realistic attention score distributions. Experiments on real-world workloads demonstrate that LiteKV reduces memory usage by up to 75% compared to state-of-the-art memory-unconstrained baselines, while preserving comparable generation quality and 95% of their decoding efficiency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper18
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
- SnapKV: LLM Knows What You are Looking for Before GenerationYuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh 等NeurIPS 2024 · 被引用 1,019 次
- H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language ModelsZhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen 等NeurIPS 2023 · 被引用 1,003 次
- Compressive Transformers for Long-Range Sequence ModellingJack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Chloe Hillier 等ICLR 2020 · 被引用 833 次
- KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache QuantizationColeman Hooper, Sehoon Kim, Hiva Mohammadzadeh, Michael W. Mahoney 等NeurIPS 2024 · 被引用 738 次
相关 Paper
- Kelle: Co-design KV Caching and eDRAM for Efficient LLM Serving in Edge ComputingTianhua Xia, Sai Qian ZhangMICRO 2025 · 被引用 2 次
- Efficient Low Rank Attention for Long-Context Inference in Large Language ModelsTenghui Li, Guoxu Zhou, Xuyang Zhao, Yuning Qiu 等NeurIPS 2025 · 被引用 4 次
- KeepKV: Achieving Periodic Lossless KV Cache Compression for Efficient LLM InferenceYuxuan Tian, Zihan Wang, Yebo Peng, Aomufei Yuan 等AAAI 2026
- Value-Guided KV Compression for LLMs via Approximated CUR DecompositionAyan Sengupta, Siddhant Chaudhary, Tanmoy ChakrabortyNeurIPS 2025 · 被引用 6 次
- Get More with LESS: Synthesizing Recurrence with KV Cache Compression for Efficient LLM InferenceHarry Dong, Xinyu Yang, Zhenyu Zhang, Zhangyang Wang 等ICML 2024 · 被引用 84 次
