Efficient Prompt Compression with Evaluator Heads for Long-Context Transformer Inference
Weizhi Fei, Xueyan Niu, Guoqing Xie, Yingqing Liu, Bo Bai, Wei Han
摘要
Although applications involving long-context inputs are crucial for the effective utilization of large language models (LLMs), they also result in increased computational costs and reduced performance. To address this challenge, we propose an efficient, training-free prompt compression method that retains key information within compressed prompts. We identify specific attention heads in transformer-based LLMs, which we designate as evaluator heads, that are capable of selecting tokens in long inputs that are most significant for inference. Building on this discovery, we develop EHPC, an Evaluator Head-based Prompt Compression method, which enables LLMs to rapidly "skim through" input prompts by leveraging only the first few layers with evaluator heads during the pre-filling stage, subsequently passing only the important tokens to the model for inference. EHPC achieves state-of-the-art results across two mainstream benchmarks: prompt compression and long-context inference acceleration. Consequently, it effectively reduces the complexity and costs associated with commercial API calls. We further demonstrate that EHPC attains competitive results compared to key-value cache-based acceleration methods, thereby highlighting its potential to enhance the efficiency of LLMs for long-context tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- DeltaEvolve: Accelerating Scientific Discovery through Momentum-Driven EvolutionJiachen Jiang, Tianyu Ding, Zhihui ZhuICML 2026 · 被引用 17 次
- Gated Differentiable Working Memory for Long-Context Language ModelingLingrui Mei, Shenghua Liu, Yiwei Wang, Yuyao Ge 等ACL 2026 · 被引用 4 次
- More with Less: An Empirical Study of Turn-Control Strategies for Efficient Coding AgentsPengfei Gao, Chao PengICSE 2026
它引用的顶会 Paper15
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
- SnapKV: LLM Knows What You are Looking for Before GenerationYuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh 等NeurIPS 2024 · 被引用 1,019 次
- DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model ServingYinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu 等OSDI 2024 · 被引用 646 次
- Learning to Compress Prompts with Gist TokensJesse Mu, Xiang Li, Noah D. GoodmanNeurIPS 2023 · 被引用 488 次
- LongCoder: A Long-Range Pre-trained Language Model for Code CompletionDaya Guo, Canwen Xu, Nan Duan, Jian Yin 等ICML 2023 · 被引用 150 次
相关 Paper
- ClusterAttn: KV Cache Compression under Intrinsic Attention ClusteringMinwei Zhang, Haifeng Sun, Jingyu Wang, Shaolong Li 等ACL 2025 · 被引用 5 次
- Prompt Compression with Context-Aware Sentence Encoding for Fast and Improved LLM InferenceBarys Liskavets, Maxim Ushakov, Shuvendu Roy, Mark Klibanov 等AAAI 2025 · 被引用 41 次
- RefreshKV: Updating Small KV Cache During Long-form GenerationFangyuan Xu, Tanya Goyal, Eunsol ChoiACL 2025 · 被引用 6 次
- HeteroCache: A Dynamic Retrieval Approach to Heterogeneous KV Cache Compression for Long-Context LLM InferenceZhiyuan Shi, Qibo Qiu, Feng Xue, Zhonglin Jiang 等ACL 2026 · 被引用 1 次
- RazorAttention: Efficient KV Cache Compression Through Retrieval HeadsHanlin Tang, Yang Lin, Jing Lin, Qingsen Han 等ICLR 2025
