D2O: Dynamic Discriminative Operations for Efficient Long-Context Inference of Large Language Models
Zhongwei Wan, Xinjian Wu, Yu Zhang, Yi Xin, Chaofan Tao, Zhihong Zhu, Xin Wang, Siqi Luo, Jing Xiong, Longyue Wang, Mi Zhang
Abstract
Generative inference in Large Language Models (LLMs) is impeded by the growing memory demands of Key-Value (KV) cache, especially for longer sequences. Traditional KV cache eviction strategies, which discard less critical KV pairs based on attention scores, often degrade generation quality, leading to issues such as context loss or hallucinations. In this work, we introduce Dynamic Discriminative Operations (D 2 O), a KV cache compression method that optimizes KV cache size dynamically and discriminatively at two levels without fine-tuning, while preserving essential context. At layer level, D 2 O leverages the varying densities of attention weights between shallow and deep layers to dynamically determine which layers should avoid excessive eviction via a novel dynamic allocation strategy to minimize information loss. At token level, D 2 O incorporates a compensation mechanism that maintains a similarity threshold to re-discriminate the importance of currently discarded tokens, determining whether they should be recalled and merged with similar tokens. We conduct experiments on various benchmarks and LLM architectures. Our results show that D 2 O not only achieves significant memory savings and enhances inference throughput by more than 3× but also maintains high-quality long-text generation. * Project leader. † Equal contribution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- Neural Attention SearchDifan Deng, Marius LindauerNeurIPS 2025 · 431 citations
- SRPO: Enhancing Multimodal LLM Reasoning via Reflection-Aware Reinforcement LearningZhongwei Wan, Zhihao Dou, Che Liu, Yu Zhang et al.NeurIPS 2025 · 63 citations
- Fast KV Compaction via Attention MatchingAdam Zweiger, Xinghong Fu, Han Guo, Yoon KimICML 2026 · 14 citations
- FreqKV: Key-Value Compression in Frequency Domain for Context Window ExtensionJushi Kai, Yixuan Wang, Boyi Zeng, Haoli Bai et al.ICLR 2026 · 7 citations
- Homogeneous Keys, Heterogeneous Values: Exploiting Local KV Cache Asymmetry for Long-Context LLMsWanyun Cui, Mingwei XuNeurIPS 2025 · 7 citations
Builds on20
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- SnapKV: LLM Knows What You are Looking for Before GenerationYuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh et al.NeurIPS 2024 · 1,019 citations
- H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language ModelsZhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen et al.NeurIPS 2023 · 1,003 citations
- Compressive Transformers for Long-Range Sequence ModellingJack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Chloe Hillier et al.ICLR 2020 · 833 citations
Related papers
- Dynamic Memory Compression: Retrofitting LLMs for Accelerated InferencePiotr Nawrot, Adrian Lancucki, Marcin Chochowski, David Tarjan et al.ICML 2024 · 106 citations
- DiffKV: Differentiated Memory Management for Large Language Models with Parallel KV CompactionYanqi Zhang, Yuwei Hu, Runyuan Zhao, John C. S. Lui et al.SOSP 2025
- Model Tells You What to Discard: Adaptive KV Cache Compression for LLMsSuyu Ge, Yunan Zhang, Liyuan Liu, Minjia Zhang et al.ICLR 2024 · 432 citations
- HitKV: Activation Frequency Knows Which Tokens Are ImportantSanle Zhao, Yujuan Tan, Jing Yu, Zhuoxin Bai et al.AAAI 2026
- RefreshKV: Updating Small KV Cache During Long-form GenerationFangyuan Xu, Tanya Goyal, Eunsol ChoiACL 2025 · 6 citations
