Lune

INFOCOM2026顶会

Memory-Efficient KV Cache Optimization for Large Language Model Inference at the Edge

Chi Zhang, Haisheng Tan, Haotian Pan, Yang Xu, Haohua Du, Li Zhang, Xiaoming Fu

2026年份
1被引次数

摘要

The deployment of large language models (LLMs) on edge devices offers notable advantages in privacy preservation, low-latency response, and operational autonomy, but faces significant challenges due to the memory-intensive key-value (KV) cache required for autoregressive decoding. Sparse attention methods alleviate computational costs but typically fail to preserve the generation quality and inference efficiency under memory constraints, which becomes particularly acute on edge platforms. In this paper, we first formalize and prove the NP-hardness of the KV cache replacement problem under top-k sparse attention, where k is the number of pages involved in each attention calculation. Then, we propose LiteKV, a lightweight memory-efficient online KV cache replacement algorithm tailored for edge-based LLM inference. LiteKV introduces a (1 + ϵ)-relaxation strategy that tolerates minimal precision loss in attention scores, allowing significant improvements in cache utilization. Moreover, LiteKV is orthogonal to existing sparse attention and quantization methods, enabling scalable, high-performance LLM inference under stringent memory budgets. We provide theoretical guarantees on LiteKV’s competitive ratio under realistic attention score distributions. Experiments on real-world workloads demonstrate that LiteKV reduces memory usage by up to 75% compared to state-of-the-art memory-unconstrained baselines, while preserving comparable generation quality and 95% of their decoding efficiency.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper18

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖