Pretraining Context Compressor for Large Language Models with Embedding-Based Memory
Yuhong Dai, Jianxun Lian, Yitian Huang, Wei Zhang, Mingyang Zhou, Mingqi Wu, Xing Xie, Hao Liao
摘要
Efficient processing of long contexts in large language models (LLMs) is essential for realworld applications like retrieval-augmented generation and in-context learning, especially in resource-constrained environments such as edge computing. This paper explores the embedding-based context compression to reduce inference costs while preserving the downstream LLM configurations. We propose a decoupled compressor-LLM framework, pretrained on text reconstruction and completion tasks, designed to effectively preserve essential contextual information within condensed embedding representations. Our extensive experiments investigate pretraining, model configurations, compression rates, efficiency across tasks, and adaptability to various LLMs. Results demonstrate that our approach outperforms competitive baselines in three domains and across eight datasets while being adaptable to different downstream LLMs. We find that thorough pretraining and carefully selected compression rates, such as 4x and 16x, enable a lightweight compressor to achieve a good balance between accuracy and speed. These findings underscore the potential of embeddingbased compression to enhance LLM efficiency and motivate further research in this area.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- GMSA: Enhancing Context Compression via Group Merging and Layer Semantic AlignmentJiwei Tang, Zhicheng Zhang, Shunlong Wu, Jingheng Ye 等ACL 2026 · 被引用 24 次
- COMI: Coarse-to-fine Context Compression via Marginal Information GainJiwei Tang, Shilei Liu, Zhicheng Zhang, Yujin Yuan 等ICLR 2026 · 被引用 17 次
- Read As Human: Compressing Context via Parallelizable Close Reading and SkimmingJiwei Tang, Shilei Liu, Zhicheng Zhang, Qingsong Lv 等ACL 2026 · 被引用 10 次
- Do LLMs Forget What They Should? Evaluating In-Context Forgetting in Large Language ModelsYuli Qian, Zechuan Yang, Wenbiao Ding, Hongzhi Li 等ICLR 2026
- Frozen LLMs are Native Decoders for High-Norm Semantic VectorsYunsheng Zeng, Yongmei TanACL 2026
它引用的顶会 Paper13
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse AttentionHuiqiang Jiang, Yucheng Li, Chengruidong Zhang, Qianhui Wu 等NeurIPS 2024 · 被引用 479 次
- MemoryBank: Enhancing Large Language Models with Long-Term MemoryWanjun Zhong, Lianghong Guo, Qiqi Gao, He Ye 等AAAI 2024 · 被引用 394 次
- LongRoPE: Extending LLM Context Window Beyond 2 Million TokensYiran Ding, Li Lyna Zhang, Chengruidong Zhang, Yuanyuan Xu 等ICML 2024 · 被引用 316 次
- LongLoRA: Efficient Fine-tuning of Long-Context Large Language ModelsYukang Chen, Shengju Qian, Haotian Tang, Xin Lai 等ICLR 2024 · 被引用 254 次
相关 Paper
- C2KV: Compressed and Composable KV Cache Reuse for Efficient LLM InferenceChuheng Du, Junyi Chen, Hanlin Tang, Kan Liu 等KDD 2026 · 被引用 3 次
- LLoCO: Learning Long Contexts OfflineSijun Tan, Xiuyu Li, Shishir G. Patil, Ziyang Wu 等EMNLP 2024 · 被引用 3 次
- Provence: efficient and robust context pruning for retrieval-augmented generationNadezhda Chirkova, Thibault Formal, Vassilina Nikoulina, Stéphane ClinchantICLR 2025 · 被引用 2 次
- Learning to Compress: Unlocking the Potential of Large Language Models for Text RepresentationYeqin Zhang, Yizheng Zhao, Chen Hu, Binxing Jiao 等AAAI 2026 · 被引用 2 次
- Compressing Context to Enhance Inference Efficiency of Large Language ModelsYucheng Li, Bo Dong, Frank Guerin, Chenghua LinEMNLP 2023 · 被引用 54 次
