Block-Attention for Efficient Prefilling
Dongyang Ma, Yan Wang, Tian Lan
Abstract
We introduce Block-attention, an attention mechanism designed to address the increased inference latency and cost in Retrieval-Augmented Generation (RAG) scenarios. Traditional approaches often encode the entire context in an autoregressive manner. Instead, Block-attention divides retrieved documents into discrete blocks, with each block independently calculating key-value (KV) states except for the final block. In RAG scenarios, by defining each passage as a block, Block-attention enables us to reuse the KV states of passages that have been seen before, thereby significantly reducing the latency and the computation overhead during inference. The implementation of Block-attention involves block segmentation, position re-encoding, and fine-tuning the LLM to adapt to the Block-attention mechanism. Experiments on 11 diverse benchmarks, including RAG, ICL, and general domains, demonstrate that after block fine-tuning, the Block-attention model not only achieves performance comparable to that of fullattention models, but can also seamlessly switch between the block and full attention modes without any performance loss. Notably, Block-attention significantly reduces the time to first token (TTFT) and floating point operations (FLOPs) to a very low level. It only takes 45 ms to output the first token for an input sequence with a total length of 32K. Compared to the full-attention models, the TTFT and corresponding FLOPs are reduced by 98.7% and 99.8%, respectively. Additionally, in Appendix A, we elaborate on how Block-attention is applied in Game AI scenario and the substantial potential benefits it entails. We strongly suggest researchers in the gaming field not to overlook this section. 1 * Equal Contribution † Corresponding Author 1 Codes, datasets and model weights have been publicly available at https://github.com/ TemporaryLoRA/Block-attention .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 63a80b15-14aa-42dd-9aa2-958f4dfda866Cited by top-tier papers9
- VMoBA: Mixture-of-Block Attention for Video Diffusion ModelsJianzong Wu, Liang Hou, Haotian Yang, Ye Tian et al.ICLR 2026 · 36 citations
- Graph-KV: Breaking Sequence via Injecting Structural Biases into Large Language ModelsHaoyu Wang, Peihao Wang, Mufei Li, Shikun Liu et al.NeurIPS 2025 · 6 citations
- From Prefix Cache to Fusion RAG Cache: Accelerating LLM Inference in Retrieval-Augmented GenerationJiahao Wang, Weiyu Xie, Mingxing Zhang, Boxin Zhang et al.SIGMOD 2026 · 4 citations
- C2KV: Compressed and Composable KV Cache Reuse for Efficient LLM InferenceChuheng Du, Junyi Chen, Hanlin Tang, Kan Liu et al.KDD 2026 · 3 citations
- TurboRAG: Accelerating Retrieval-Augmented Generation with Precomputed KV Caches for Chunked TextSongshuo Lu, Hua Wang, Yutian Rong, Zhi Chen et al.EMNLP 2025 · 2 citations
Builds on20
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu et al.NeurIPS 2023 · 5,989 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionAkari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil et al.ICLR 2024 · 1,798 citations
- SGLang: Efficient Execution of Structured Language Model ProgramsLianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun et al.NeurIPS 2024 · 1,586 citations
Related papers
- LazyAttention: Efficient Retrieval-Augmented Generation with Deferred Positional EncodingHaocheng Xia, Mihir Pamnani, Hanxi Fang, Supawit Chockchowwat et al.ICML 2026
- Accelerating Inference of Retrieval-Augmented Generation via Sparse Context SelectionYun Zhu, Jia-Chen Gu, Caitlin Sikora, Ho Ko et al.ICLR 2025
- Sparse Attention Across Multiple-Context KV CacheZiyi Cao, Qingyi Si, Jingbin Zhang, Bingquan LiuAAAI 2026 · 3 citations
- ProphetKV: User-Query-Driven Selective Recomputation for Efficient KV Cache Reuse in Retrieval-Augmented GenerationShihao Wang, Jiahao Chen, Yanqi Pan, Hao Huang et al.ICML 2026 · 4 citations
- DepCache: A KV Cache Management Framework for GraphRAG with Dependency AttentionHao Yuan, Xin Ai, Qiange Wang, Peizheng Li et al.SIGMOD 2026 · 2 citations
