Leveraging Attention to Effectively Compress Prompts for Long-Context LLMs
Yunlong Zhao, Haoran Wu, Bo Xu
Abstract
Prompt compression is increasingly studied for its potential to reduce computational costs and alleviate the burden on language models when processing lengthy prompts. Prior research has assessed token retention and removal by calculating information entropy. However, prompt compression encounters two significant challenges: (1) Information entropy, while widely used, may not be the optimal compression metric; and (2) The semantic significance of tokens is context-dependent, which renders independent token retention decisions inadequate.
We posit that the solution to these challenges lies in the intrinsic mechanism of language models. Large language models (LLMs) exhibit robust contextual processing capabilities, with recent studies on their internal dynamics revealing that the attention mechanism plays a crucial role in modeling how LLMs leverage long contexts. Building on this insight, we introduce AttnComp, a novel approach that exploits the attention mechanism within language models to guide prompt compression. Our method employs causal cross-attention from the query to the context to evaluate the significance of each token, and we develop a graph-based algorithm to efficiently cluster tokens into semantic units, thus mitigating the issue of independent dependencies.
We conduct experiments on datasets for retrieval-augmented generation and multiple long tasks involving single or multi-document QA. Our proposed method, AttnComp, outperforms previous baselines and validates the contributions of our components through analytical experiments. Compared to other methods that use a causal LM for prompt compression, our approach results in shorter latency and improved performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- GMSA: Enhancing Context Compression via Group Merging and Layer Semantic AlignmentJiwei Tang, Zhicheng Zhang, Shunlong Wu, Jingheng Ye et al.ACL 2026 · 24 citations
- COMI: Coarse-to-fine Context Compression via Marginal Information GainJiwei Tang, Shilei Liu, Zhicheng Zhang, Yujin Yuan et al.ICLR 2026 · 17 citations
- Revitalizing Black-Box Interpretability: Actionable Interpretability for LLMs via Proxy ModelsJunhao Liu, Haonan Yu, Zhenyu Yan, Xin ZhangACL 2026 · 2 citations
- HiGoE: Hierarchical Graph of Evidence to Enhance Retrieval-Augmented Generation for Long-context SummarizationLong Yuan, Kaiwen Tian, Zi Chen, Bolong Zheng et al.ACL 2026
- CGMIS: Concept-Graph Based Multi-Hop Instructions Synthesis for Enhancing Long-Context ReasoningZechen Sun, Zecheng Tang, Juntao Li, Wenpeng Hu et al.AAAI 2026
Builds on21
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris et al.UIST 2023 · 1,882 citations
- Learning to Compress Prompts with Gist TokensJesse Mu, Xiang Li, Noah D. GoodmanNeurIPS 2023 · 488 citations
- A Survey on In-context LearningQingxiu Dong, Lei Li, Damai Dai, Ce Zheng et al.EMNLP 2024 · 479 citations
- RECOMP: Improving Retrieval-Augmented LMs with Context Compression and Selective AugmentationFangyuan Xu, Weijia Shi, Eunsol ChoiICLR 2024 · 260 citations
Related papers
- DAC: A Dynamic Attention-aware Approach for Task-Agnostic Prompt CompressionYi Zhao, Zuchao Li, Hai Zhao, Baoyuan Qi et al.ACL 2025 · 7 citations
- Prompt Compression with Context-Aware Sentence Encoding for Fast and Improved LLM InferenceBarys Liskavets, Maxim Ushakov, Shuvendu Roy, Mark Klibanov et al.AAAI 2025 · 41 citations
- LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt CompressionHuiqiang Jiang, Qianhui Wu, Xufang Luo, Dongsheng Li et al.ACL 2024 · 59 citations
- Attn-GS: Attention-Guided Context Compression for Efficient Personalized LLMsShenglai Zeng, Tianqi Zheng, Chuan Tian, Dante Everaert et al.ACL 2026 · 1 citation
- Rethinking Token Reduction for Large Vision-Language ModelsYi Wang, Haofei Zhang, Qihan Huang, Anda Cao et al.CVPR 2026 · 1 citation
