DAC: A Dynamic Attention-aware Approach for Task-Agnostic Prompt Compression
Yi Zhao, Zuchao Li, Hai Zhao, Baoyuan Qi, Guoming Liu
摘要
Task-agnostic prompt compression leverages the redundancy in natural language to reduce computational overhead and enhance information density within prompts, especially in longcontext scenarios. Existing methods predominantly rely on information entropy as the metric to compress lexical units, aiming to achieve minimal information loss. However, these approaches overlook two critical aspects: (i) the importance of attention-critical tokens at the algorithmic level, and (ii) shifts in information entropy during the compression process. Motivated by these challenges, we propose a dynamic attention-aware approach for taskagnostic prompt compression (DAC). This approach effectively integrates entropy and attention information, dynamically sensing entropy shifts during compression to achieve fine-grained prompt compression. Extensive experiments across various domains, including LongBench, GSM8K, and BBH, show that DAC consistently yields robust and substantial improvements across a diverse range of tasks and LLMs, offering compelling evidence of its efficacy. Our code is available at https://github.com/QQQ-yi/DAC *
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Scaling LLM Speculative Decoding: Non-Autoregressive Forecasting in Large-Batch ScenariosLuohe Shi, Zuchao Li, Lefei Zhang, Baoyuan Qi 等AAAI 2026 · 被引用 1 次
- Less Is More: Elevating RAG via Performance-Driven Context CompressionZiqiang Cui, Yunpeng Weng, Xing Tang, Peiyang Liu 等ICML 2026
- End-to-End Contrastive Language-Speech Pretraining Model for Long-Form Spoken Question AnsweringJiliang Hu, Zuchao Li, Baoyuan Qi, Guoming Liu 等AAAI 2026
- Faster In-Context Learning for LLMs via N-Gram Trie Speculative DecodingJinglin Chen, Qiwei Li, Zuchao Li, Baoyuan Qi 等EMNLP 2025
它引用的顶会 Paper12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Scissorhands: Exploiting the Persistence of Importance Hypothesis for LLM KV Cache Compression at Test TimeZichang Liu, Aditya Desai, Fangshuo Liao, Weitao Wang 等NeurIPS 2023 · 被引用 557 次
- A Survey on In-context LearningQingxiu Dong, Lei Li, Damai Dai, Ce Zheng 等EMNLP 2024 · 被引用 479 次
相关 Paper
- Leveraging Attention to Effectively Compress Prompts for Long-Context LLMsYunlong Zhao, Haoran Wu, Bo XuAAAI 2025 · 被引用 10 次
- Prompt Compression with Context-Aware Sentence Encoding for Fast and Improved LLM InferenceBarys Liskavets, Maxim Ushakov, Shuvendu Roy, Mark Klibanov 等AAAI 2025 · 被引用 41 次
- GMSA: Enhancing Context Compression via Group Merging and Layer Semantic AlignmentJiwei Tang, Zhicheng Zhang, Shunlong Wu, Jingheng Ye 等ACL 2026 · 被引用 24 次
- EntroKV: Entropy-Guided Dynamic Budget Allocation for KV-Cache CompressionWenhao Gao, Haoran Cao, Yueyan Li, YongGao Xiao 等ICML 2026
- Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token CompressorsXiangchen Wang, Jinrui Zhang, Teng Wang, Haigang Zhang 等EMNLP 2025 · 被引用 2 次
