ProtoKV: Long-context Knowledges Are Already Well-Organized Before Your Query
Zhiyuan Yu, Shijian Xiao, Zhangyue Yin, Xiaoran Liu, Lekai Xing, Wenzhong Li, Nguyen Cam-Tu, Sanglu Lu
摘要
Modern Large Language Models (LLMs) face fundamental challenges in processing long text sequences due to the quadratic complexity of attention mechanisms. Key-Value (KV) cache retention strategies mitigate this issue by selectively preserving salient KV pairs for autoregressive generation. However, existing methods fail to adequately and efficiently preserve the semantic integrity of the compressed representations. In this paper, we discover a prevalent phenomenon in LLM: within the key embedding space, while most tokens exhibit similarity with their contextual neighbors (we term position-determined tokens), a small subset of special tokens serving as semantic anchors consistently show local deviation property and form one or several clusters (we term semantic-anchored tokens). Motivated by this observation, we propose ProtoKV that separately processes these two token categories for KV cache compression. Within this framework, we first construct semantic prototypes based on the inherent characteristics of the two token categories, and subsequently form clusters of semantically similar tokens as basic compression units. This approach preserves semantic integrity with high computational efficiency. Experiments on LongBench demonstrate that ProtoKV achieves 2.11% higher accuracy than state-of-the-art methods under matched memory constraints. Our code can be available in https://github.com/yyy0959/ProtoKV .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper33
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
- SnapKV: LLM Knows What You are Looking for Before GenerationYuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh 等NeurIPS 2024 · 被引用 1,019 次
- H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language ModelsZhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen 等NeurIPS 2023 · 被引用 1,003 次
相关 Paper
- ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM InferenceXiang Liu, Zhenheng Tang, Peijie Dong, Zeyu Li 等NeurIPS 2025 · 被引用 71 次
- IceCache: Memory-Efficient KV-cache Management for Long-Sequence LLMsYuzhen Mao, Qitong Wang, Martin Ester, Ke LiICLR 2026 · 被引用 6 次
- ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable CompressionGuangda Liu, Chengwei Li, Jieru Zhao, Chenqi Zhang 等DAC 2025 · 被引用 5 次
- Not All Heads Matter: A Head-Level KV Cache Compression Method with Integrated Retrieval and ReasoningYu Fu, Zefan Cai, Abedelkadir Asi, Wayne Xiong 等ICLR 2025
- Cache What Lasts: Token Retention for Memory-Bounded KV Cache in LLMsNgoc Bui, Shubham Sharma, Simran Lamba, Saumitra Mishra 等ICLR 2026 · 被引用 19 次
