QUEST: Query-Aware Sparsity for Efficient Long-Context LLM Inference
Jiaming Tang, Yilong Zhao, Kan Zhu, Guangxuan Xiao, Baris Kasikci, Song Han
摘要
As the demand for long-context large language models (LLMs) increases, models with context windows of up to 128K or 1M tokens are becoming increasingly prevalent. However, longcontext LLM inference is challenging since the inference speed decreases significantly as the sequence length grows. This slowdown is primarily caused by loading a large KV cache during self-attention. Previous works have shown that a small portion of critical tokens will dominate the attention outcomes. However, we observe the criticality of a token highly depends on the query. To this end, we propose Quest, a query-aware KV cache selection algorithm. Quest keeps track of the minimal and maximal Key values in KV cache pages and estimates the criticality of a given page using Query vectors. By only loading the Top-K critical KV cache pages for attention, Quest significantly speeds up self-attention without sacrificing accuracy. We show that Quest can achieve up to 7.03× self-attention speedup, which reduces inference latency by 2.23× while performing well on tasks with long dependencies with negligible accuracy loss. Code is available at https: //github.com/mit-han-lab/Quest .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper167
- MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse AttentionHuiqiang Jiang, Yucheng Li, Chengruidong Zhang, Qianhui Wu 等NeurIPS 2024 · 被引用 479 次
- Neural Attention SearchDifan Deng, Marius LindauerNeurIPS 2025 · 被引用 431 次
- Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse AttentionJingyang Yuan, Huazuo Gao, Damai Dai, Junyu Luo 等ACL 2025 · 被引用 334 次
- MoBA: Mixture of Block Attention for Long-Context LLMsEnzhe Lu, Zhejun Jiang, Jingyuan Liu, Yulun Du 等NeurIPS 2025 · 被引用 219 次
- RetrievalAttention: Accelerating Long-Context LLM Inference via Vector RetrievalDi Liu, Meng Chen, Baotong Lu, Huiqiang Jiang 等NeurIPS 2025 · 被引用 148 次
它引用的顶会 Paper13
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language ModelsZhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen 等NeurIPS 2023 · 被引用 1,003 次
- Compressive Transformers for Long-Range Sequence ModellingJack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Chloe Hillier 等ICLR 2020 · 被引用 833 次
相关 Paper
- TokenSelect: Efficient Long-Context Inference and Length Extrapolation for LLMs via Dynamic Token-Level KV Cache SelectionWei Wu, Zhuoshi Pan, Kun Fu, Chao Wang 等EMNLP 2025 · 被引用 2 次
- Vegas: Self-Speculative Decoding with Verification-Guided Sparse AttentionYikang Yue, Yuqi Xue, Jian HuangICML 2026 · 被引用 2 次
- KVzip: Query-Agnostic KV Cache Compression with Context ReconstructionJang-Hyun Kim, Jinuk Kim, Sangwoo Kwon, Jae W. Lee 等NeurIPS 2025 · 被引用 103 次
- Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token SelectionDongwon Jo, Beomseok Kang, Jiwon Song, jae-joon kimICML 2026 · 被引用 1 次
- ThinK: Thinner Key Cache by Query-Driven PruningYuhui Xu, Zhanming Jie, Hanze Dong, Lei Wang 等ICLR 2025
