Self-Indexing KVCache: Predicting Sparse Attention from Compressed Keys
Xu Yang, Jiapeng Zhang, Dongyang Zhao, Guo Chen, Zhuo Tang
Abstract
The KV cache in self-attention has emerged as a major bottleneck in long-context and large-batch inference for LLMs. Existing approaches often treat sparsity prediction and compression as separate modules—relying on auxiliary index structures to select relevant tokens, and on complex quantization schemes to reduce memory usage. This fragmented design introduces redundant overhead and limits scalability.
In this paper, we propose a novel paradigm: treating the compressed key representation not merely as storage, but as a self-indexing structure that directly enables efficient sparse attention. By designing a sign-based 1-bit vector quantization (VQ) scheme, our method unifies compression and retrieval in a single, hardware-friendly format. This approach eliminates the need for external indices or learning-based predictors, offering a lightweight yet robust solution for memory-constrained inference.
All components are designed to be hardware-efficient and easy to implement. By implementing custom CUDA kernels, our method integrates seamlessly with FlashAttention, minimizing additional runtime and memory overhead. Experimental results demonstrate that our approach delivers both effectiveness and efficiency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 936b258a-54f8-4594-9f41-5dfd4a71ee1fBuilds on14
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu et al.ICML 2023 · 1,493 citations
- SnapKV: LLM Knows What You are Looking for Before GenerationYuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh et al.NeurIPS 2024 · 1,019 citations
- KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache QuantizationColeman Hooper, Sehoon Kim, Hiva Mohammadzadeh, Michael W. Mahoney et al.NeurIPS 2024 · 738 citations
Related papers
- MUSTAFAR: Promoting Unstructured Sparsity for KV Cache Pruning in LLM InferenceDonghyeon Joo, Helya Hosseini, Ramyad Hadidi, Bahar AsgariNeurIPS 2025 · 12 citations
- JanusQuant: Accurate and Efficient 2-bit KV Cache Quantization for Long-Context InferenceChengyu Sun, Yaqi Xia, Hulin Wang, Donglin Yang et al.PPoPP 2026 · 1 citation
- MILLION: MasterIng Long-Context LLM Inference Via Outlier-Immunized KV Product QuaNtizationZongwu Wang, Peng Xu, Fangxin Liu, Yiwei Hu et al.DAC 2025 · 6 citations
- ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token IdentificationYefei He, Luoming Zhang, Weijia Wu, Jing Liu et al.NeurIPS 2024 · 100 citations
- RaBitQCache: Rotated Binary Quantization for KVCache in Long Context LLM InferenceWenhao Li, Jinhao Dong, Hailin Zhang, Wenhang Shi et al.ICML 2026
