vAttention: Verified Sparse Attention via Sampling
Aditya Desai, Kumar Krishna Agrawal, Shuo Yang, Alejandro Cuadron, Luis Gaspar Schroeder, Matei Zaharia, Joseph E. Gonzalez, Ion Stoica
摘要
State-of-the-art sparse attention methods for reducing decoding latency fall into two main categories: approximate top- (and its extension, top-) and recently introduced sampling-based estimation. However, these approaches are fundamentally limited in their ability to approximate full attention: they fail to provide consistent approximations across heads and query vectors and, most critically, lack guarantees on approximation quality, limiting their practical deployment. We observe that top- and random sampling are complementary: top- performs well when attention scores are dominated by a few tokens, whereas random sampling provides better estimates when attention scores are relatively uniform. Building on this insight and leveraging the statistical guarantees of sampling, we introduce vAttention, the first practical sparse attention mechanism with user-specified guarantees on approximation accuracy. These guarantees make vAttention a compelling step toward practical, reliable deployment of sparse attention at scale. By unifying top-k and sampling, vAttention outperforms both individually, delivering a superior quality–efficiency trade-off. Our experiments show that vAttention significantly improves the quality of sparse attention (e.g., 4.5 percentage points for Llama-3.1-8B-Inst and Deepseek-R1-Distill-Llama-8B on RULER-HARD ), and effectively bridges the gap between full and sparse attention (e.g., across datasets, it matches full model quality at 10x–20x sparsity). We also demonstrate that it can be deployed in long-generation scenarios to achieve fast decoding without compromising model quality (e.g., vAttention achieves full model quality on AIME2024 at 10% sparsity with up to 32K token generations).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
- SnapKV: LLM Knows What You are Looking for Before GenerationYuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh 等NeurIPS 2024 · 被引用 1,019 次
- H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language ModelsZhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen 等NeurIPS 2023 · 被引用 1,003 次
- Scissorhands: Exploiting the Persistence of Importance Hypothesis for LLM KV Cache Compression at Test TimeZichang Liu, Aditya Desai, Fangshuo Liao, Weitao Wang 等NeurIPS 2023 · 被引用 557 次
相关 Paper
- MagicPIG: LSH Sampling for Efficient LLM GenerationZhuoming Chen, Ranajoy Sadhukhan, Zihao Ye, Yang Zhou 等ICLR 2025 · 被引用 1 次
- Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token SelectionDongwon Jo, Beomseok Kang, Jiwon Song, jae-joon kimICML 2026 · 被引用 1 次
- SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM PrefillingXiaodong Ji, Hailin Zhang, Fangcheng Fu, Bin CuiICML 2026 · 被引用 3 次
- Tactic: Adaptive Sparse Attention with Clustering and Distribution Fitting for Long-Context LLMsKan Zhu, Tian Tang, Qinyu Xu, Zhan Jin 等ICLR 2026 · 被引用 25 次
- RaBitQCache: Rotated Binary Quantization for KVCache in Long Context LLM InferenceWenhao Li, Jinhao Dong, Hailin Zhang, Wenhang Shi 等ICML 2026
