Streaming Attention Approximation via Discrepancy Theory
Ekaterina Kochetkova, Kshiteej Sheth, Insu Han, Amir Zandieh, Michael Kapralov
Abstract
Large language models (LLMs) have achieved impressive success, but their high memory requirements present challenges for long-context token generation. In this paper we study the streaming complexity of attention approximation, a key computational primitive underlying token generation. Our main contribution is BalanceKV, a streaming algorithm for -approximating attention computations based on geometric process for selecting a balanced collection of Key and Value tokens as per Banaszczyk's vector balancing theory. We complement our algorithm with space lower bounds for streaming attention computation. Besides strong theoretical guarantees, BalanceKV exhibits empirically validated performance improvements over existing methods, both for attention approximation and end-to-end performance on various long context benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0598c2a1-93e0-41c7-855d-aef79317fdd7Cited by top-tier papers2
- Accurate KV Cache Eviction via Anchor Direction Projection for Efficient LLM InferenceZijie Geng, Jie Wang, Ziqi Liu, Feng Ju et al.NeurIPS 2025 · 6 citations
- WildCat: Near-Linear Attention in Theory and PracticeTobias Schröder, Lester MackeyICML 2026 · 3 citations
Builds on32
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- SnapKV: LLM Knows What You are Looking for Before GenerationYuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh et al.NeurIPS 2024 · 1,019 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
Related papers
- Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token SelectionDongwon Jo, Beomseok Kang, Jiwon Song, jae-joon kimICML 2026 · 1 citation
- TokenSelect: Efficient Long-Context Inference and Length Extrapolation for LLMs via Dynamic Token-Level KV Cache SelectionWei Wu, Zhuoshi Pan, Kun Fu, Chao Wang et al.EMNLP 2025 · 2 citations
- ArkVale: Efficient Generative LLM Inference with Recallable Key-Value EvictionRenze Chen, Zhuofeng Wang, Beiquan Cao, Tong Wu et al.NeurIPS 2024 · 56 citations
- Latent-Condensed Transformer for Efficient Long Context ModelingZeng You, Yaofo Chen, Qiuwu Chen, Ying Sun et al.ACL 2026
- Token Sample Complexity of AttentionLéa Bohbot, Cyril Letrouit, Gabriel Peyré, François-Xavier VialardICML 2026 · 1 citation
