OmniKV: Dynamic Context Selection for Efficient Long-Context LLMs
Jitai Hao, Yuke Zhu, Tian Wang, Jun Yu, Xin Xin, Bo Zheng, Zhaochun Ren, Sheng Guo
摘要
During the inference phase of Large Language Models (LLMs) with long context, a substantial portion of GPU memory is allocated to the KV cache, with memory usage increasing as the sequence length grows. To mitigate the GPU memory footprint associate with KV cache, some previous studies have discarded less important tokens based on the sparsity identified in attention scores in long context scenarios. However, we argue that attention scores cannot indicate the future importance of tokens in subsequent generation iterations, because attention scores are calculated based on current hidden states. Therefore, we propose Om-niKV, a token-dropping-free and training-free inference method, which achieves a 1.68x speedup without any loss in performance. It is well-suited for offloading, significantly reducing KV cache memory usage by up to 75% with it. The core innovative insight of OmniKV is: Within a single generation iteration, there is a high degree of similarity in the important tokens identified across consecutive layers. Extensive experiments demonstrate that OmniKV achieves state-of-the-art performance across multiple benchmarks, with particularly advantages in chainof-thoughts scenarios. OmniKV extends the maximum context length supported by a single A100 for Llama-3-8B from 128K to 450K. Our code is available at https://github.com/antgroup/OmniKV.git
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- ThinKV: Thought-Adaptive KV Cache Compression for Efficient Reasoning ModelsAkshat Ramachandran, Marina Neseem, Charbel Sakr, Rangharajan Venkatesan 等ICLR 2026 · 被引用 19 次
- Uni-X: Mitigating Modality Conflict with a Two-End-Separated Architecture for Unified Multimodal ModelsJitai Hao, Hao Liu, Xinyan Xiao, Qiang Huang 等ICLR 2026 · 被引用 18 次
- A Token is Worth over 1, 000 Tokens: Efficient Knowledge Distillation through Low-Rank CloneJitai Hao, Qiang Huang, Hao Liu, Xinyan Xiao 等NeurIPS 2025 · 被引用 17 次
- Sparse Attention Adaptation for Long ReasoningYizhao Gao, Shuming Guo, Shijie Cao, Yuqing Xia 等ICLR 2026 · 被引用 17 次
- CacheSlide: Unlocking Cross Position-Aware KV Cache Reuse for Accelerating LLM ServingYang Liu, Yunfei Gu, Liqiang Zhang, Chentao Wu 等FAST 2026 · 被引用 14 次
它引用的顶会 Paper26
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu 等ICML 2023 · 被引用 1,493 次
- SnapKV: LLM Knows What You are Looking for Before GenerationYuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh 等NeurIPS 2024 · 被引用 1,019 次
相关 Paper
- TokenSelect: Efficient Long-Context Inference and Length Extrapolation for LLMs via Dynamic Token-Level KV Cache SelectionWei Wu, Zhuoshi Pan, Kun Fu, Chao Wang 等EMNLP 2025 · 被引用 2 次
- KeyDiff: Key Similarity-Based KV Cache Eviction for Long-Context LLM Inference in Resource-Constrained EnvironmentsJunyoung Park, Dalton Jones, Matthew J. Morse, Raghavv Goel 等NeurIPS 2025 · 被引用 47 次
- SparK: Query-Aware Unstructured Sparsity with Recoverable KV Cache Channel PruningHuanxuan Liao, Yixing Xu, Shizhu He, Guanchen Li 等AAAI 2026 · 被引用 3 次
- KVzip: Query-Agnostic KV Cache Compression with Context ReconstructionJang-Hyun Kim, Jinuk Kim, Sangwoo Kwon, Jae W. Lee 等NeurIPS 2025 · 被引用 103 次
- RefreshKV: Updating Small KV Cache During Long-form GenerationFangyuan Xu, Tanya Goyal, Eunsol ChoiACL 2025 · 被引用 6 次
