Compressed Context Memory for Online Language Model Interaction
Jang-Hyun Kim, Junyoung Yeom, Sangdoo Yun, Hyun Oh Song
摘要
This paper presents a context key/value compression method for Transformer language models in online scenarios, where the context continually expands. As the context lengthens, the attention process demands increasing memory and computations, which in turn reduces the throughput of the language model. To address this challenge, we propose a compressed context memory system that continually compresses the accumulating attention key/value pairs into a compact memory space, facilitating language model inference in a limited memory space of computing environments. Our compression process involves integrating a lightweight conditional LoRA into the language model's forward pass during inference, without the need for fine-tuning the model's entire set of weights. We achieve efficient training by modeling the recursive compression process as a single parallelized forward computation. Through evaluations on conversation, personalization, and multi-task learning, we demonstrate that our approach achieves the performance level of a full context model with smaller context memory size. We further demonstrate the applicability of our approach in a streaming setting with an unlimited context length, outperforming the sliding window approach. Codes are available at https://github.com/snu-mllab/context-memory.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- KVzip: Query-Agnostic KV Cache Compression with Context ReconstructionJang-Hyun Kim, Jinuk Kim, Sangwoo Kwon, Jae W. Lee 等NeurIPS 2025 · 被引用 103 次
- R-KV: Redundancy-aware KV Cache Compression for Reasoning ModelsZefan Cai, Wen Xiao, Hanshi Sun, Cheng Luo 等NeurIPS 2025 · 被引用 50 次
- Online Adaptation of Language Models with a Memory of Amortized ContextsJihoon Tack, Jaehyung Kim, Eric Mitchell, Jinwoo Shin 等NeurIPS 2024 · 被引用 46 次
- Multi-Head Low-Rank AttentionSongtao Liu, Hongwu Peng, Zhiwei Zhang, Zhengyu Chen 等ICLR 2026 · 被引用 18 次
- Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data FormatChao Fang, Man Shi, Robin Geens, Arne Symons 等HPCA 2025 · 被引用 15 次
它引用的顶会 Paper16
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
- Compressive Transformers for Long-Range Sequence ModellingJack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Chloe Hillier 等ICLR 2020 · 被引用 833 次
- Learning to Compress Prompts with Gist TokensJesse Mu, Xiang Li, Noah D. GoodmanNeurIPS 2023 · 被引用 488 次
相关 Paper
- Dynamic Memory Compression: Retrofitting LLMs for Accelerated InferencePiotr Nawrot, Adrian Lancucki, Marcin Chochowski, David Tarjan 等ICML 2024 · 被引用 106 次
- LLoCO: Learning Long Contexts OfflineSijun Tan, Xiuyu Li, Shishir G. Patil, Ziyang Wu 等EMNLP 2024 · 被引用 3 次
- Dodo: Dynamic Contextual Compression for Decoder-only LMsGuanghui Qin, Corby Rosset, Ethan C. Chau, Nikhil Rao 等ACL 2024
- GradMem: Learning to Write Context into Memory with Test-Time Gradient DescentYuri Kuratov, Matvey Kairov, Aydar Bulatov, Ivan Rodkin 等ICML 2026 · 被引用 3 次
- Layer-Condensed KV Cache for Efficient Inference of Large Language ModelsHaoyi Wu, Kewei TuACL 2024
