Compressing Context to Enhance Inference Efficiency of Large Language Models
Yucheng Li, Bo Dong, Frank Guerin, Chenghua Lin
Abstract
Large language models (LLMs) achieved remarkable performance across various tasks. However, they face challenges in managing long documents and extended conversations, due to significantly increased computational requirements, both in memory and inference time, and potential context truncation when the input exceeds the LLM's fixed context length. This paper proposes a method called Selective Context that enhances the inference efficiency of LLMs by identifying and pruning redundancy in the input context to make the input more compact. We test our approach using common data sources requiring long context processing: arXiv papers, news articles, and long conversations, on tasks of summarisation, question answering, and response generation. Experimental results show that Selective Context significantly reduces memory cost and decreases generation latency while maintaining comparable performance compared to that achieved when full context is used. Specifically, we achieve a 50% reduction in context cost, resulting in a 36% reduction in inference memory usage and a 32% reduction in inference time, while observing only a minor drop of .023 in BERTscore and .038 in faithfulness on four downstream applications, indicating that our method strikes a good balance between efficiency and performance. Code and data are available at https://github.com/ liyucheng09/Selective_Context .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ab161c7a-5beb-43d8-8bfa-5139dc1b2ff8Cited by top-tier papers90
- Mental-LLM: Leveraging Large Language Models for Mental Health Prediction via Online Text DataXuhai Xu, Bingsheng Yao, Yuanzhe Dong, Saadia Gabriel et al.UbiComp 2024 · 281 citations
- MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory AgentHongli Yu, Tinghong Chen, Jiangtao Feng, Jiangjie Chen et al.ICLR 2026 · 231 citations
- MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon AgentsZijian Zhou, Ao Qu, Zhaoxuan Wu, Sunghwan Kim et al.ICLR 2026 · 223 citations
- C3oT: Generating Shorter Chain-of-Thought Without Compromising EffectivenessYu Kang, Xianghui Sun, Liangyu Chen, Wei ZouAAAI 2025 · 162 citations
- LightMem: Lightweight and Efficient Memory-Augmented GenerationJizhan Fang, Xinle Deng, Haoming Xu, Ziyan Jiang et al.ICLR 2026 · 162 citations
Builds on5
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Learning to Compress Prompts with Gist TokensJesse Mu, Xiang Li, Noah D. GoodmanNeurIPS 2023 · 488 citations
- Adapting Language Models to Compress ContextsAlexis Chevalier, Alexander Wettig, Anirudh Ajith, Danqi ChenEMNLP 2023 · 34 citations
Related papers
- Dynamic Context Pruning for Efficient and Interpretable Autoregressive TransformersSotiris Anagnostidis, Dario Pavllo, Luca Biggio, Lorenzo Noci et al.NeurIPS 2023 · 95 citations
- Pretraining Context Compressor for Large Language Models with Embedding-Based MemoryYuhong Dai, Jianxun Lian, Yitian Huang, Wei Zhang et al.ACL 2025
- Accelerating Inference of Retrieval-Augmented Generation via Sparse Context SelectionYun Zhu, Jia-Chen Gu, Caitlin Sikora, Ho Ko et al.ICLR 2025
- LLoCO: Learning Long Contexts OfflineSijun Tan, Xiuyu Li, Shishir G. Patil, Ziyang Wu et al.EMNLP 2024 · 3 citations
- Read As Human: Compressing Context via Parallelizable Close Reading and SkimmingJiwei Tang, Shilei Liu, Zhicheng Zhang, Qingsong Lv et al.ACL 2026 · 10 citations
