RefreshKV: Updating Small KV Cache During Long-form Generation
Fangyuan Xu, Tanya Goyal, Eunsol Choi
摘要
Generating long sequences of tokens given a long-context input is a very compute-intensive inference scenario for large language models (LLMs). One prominent inference speed-up approach is to construct a smaller key-value (KV) cache, relieving LLMs from computing attention over a long sequence of tokens. While such methods work well to generate short sequences, their performance degrades rapidly for long-form generation. Most KV compression happens once, prematurely removing tokens that can be useful later in the generation. We propose a new inference method, RefreshKV, that flexibly alternates between full context attention and attention over a subset of input tokens during generation. After each full attention step, we update the smaller KV cache based on the attention pattern over the entire input. Applying our method to off-the-shelf LLMs achieves comparable speedup to eviction-based methods while improving performance for various long-form generation tasks. Lastly, we show that continued pretraining with our inference setting brings further gains in performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- LouisKV: Efficient KV Cache Retrieval for Long Input-Output SequencesWenbo Wu, Qingyi Si, Xiurui Pan, Ye Wang 等ICLR 2026 · 被引用 5 次
- ContrastKV: Robust KV Cache Eviction via Contrastive Signal Fusion for Multi-Query GeneralizationXingchi Chen, Peiyuan Zong, Ziqiang Gao, Qing Li 等ACL 2026
它引用的顶会 Paper17
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu 等ICML 2023 · 被引用 1,493 次
- SnapKV: LLM Knows What You are Looking for Before GenerationYuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh 等NeurIPS 2024 · 被引用 1,019 次
相关 Paper
- C2KV: Compressed and Composable KV Cache Reuse for Efficient LLM InferenceChuheng Du, Junyi Chen, Hanlin Tang, Kan Liu 等KDD 2026 · 被引用 3 次
- Model Tells You What to Discard: Adaptive KV Cache Compression for LLMsSuyu Ge, Yunan Zhang, Liyuan Liu, Minjia Zhang 等ICLR 2024 · 被引用 432 次
- FreqKV: Key-Value Compression in Frequency Domain for Context Window ExtensionJushi Kai, Yixuan Wang, Boyi Zeng, Haoli Bai 等ICLR 2026 · 被引用 7 次
- Retrospective Sparse Attention for Efficient Long-Context GenerationSeonghwan Choi, Beomseok Kang, Dongwon Jo, Jae-Joon KimICLR 2026 · 被引用 4 次
- HitKV: Activation Frequency Knows Which Tokens Are ImportantSanle Zhao, Yujuan Tan, Jing Yu, Zhuoxin Bai 等AAAI 2026
