TurboRAG: Accelerating Retrieval-Augmented Generation with Precomputed KV Caches for Chunked Text
Songshuo Lu, Hua Wang, Yutian Rong, Zhi Chen, Yaohua Tang
Abstract
Current Retrieval-Augmented Generation (RAG) systems concatenate and process numerous retrieved document chunks for prefill which requires a large volume of online computation, therefore leading to significant latency in time-to-first-token (TTFT). To reduce the computation overhead as well as TTFT, we introduce TurboRAG, a hybrid offline-online paradigm that (i) pre-computes chunk-level key-value (KV) caches, (ii) stitches them together at inference time using independent-attention and reordered-RoPE techniques, and (iii) preserves answer quality without changing the model architecture. Our approach is applicable to most existing large language models and their applications without any requirement in modification of models and inference systems. Experimental results across a suite of RAG benchmarks demonstrate that TurboRAG reduces TTFT by up to 9.4x compared to the conventional RAG systems (on an average of 8.6x), but reserving comparable performance to the standard RAG systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 90b184d3-3480-4640-8ee4-f1ae50bec309Cited by top-tier papers14
- KVLink: Accelerating Large Language Models via Efficient KV Cache ReuseJingbo Yang, Bairu Hou, Wei Wei, Yujia Bao et al.NeurIPS 2025 · 83 citations
- Log-Augmented Generation: Scaling Test-Time Reasoning with Reusable ComputationPeter Baile Chen, Yi Zhang, Dan Roth, Samuel Madden et al.ICLR 2026 · 5 citations
- Efficient Many-Shot In-Context Learning with Dynamic Block-Sparse AttentionEmily Xiao, Chin-Jou Li, Yilin Zhang, Graham Neubig et al.ACL 2025 · 4 citations
- From Prefix Cache to Fusion RAG Cache: Accelerating LLM Inference in Retrieval-Augmented GenerationJiahao Wang, Weiyu Xie, Mingxing Zhang, Boxin Zhang et al.SIGMOD 2026 · 4 citations
- C2KV: Compressed and Composable KV Cache Reuse for Efficient LLM InferenceChuheng Du, Junyi Chen, Hanlin Tang, Kan Liu et al.KDD 2026 · 3 citations
Builds on14
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Improving Language Models by Retrieving from Trillions of TokensSebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai et al.ICML 2022 · 1,629 citations
- H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language ModelsZhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen et al.NeurIPS 2023 · 1,003 citations
Related papers
- SubGCache: Accelerating Graph-based RAG with Subgraph-level KV CacheQiuyu Zhu, Liang Zhang, Qianxiong Xu, Cheng Long et al.AAAI 2026 · 1 citation
- ProphetKV: User-Query-Driven Selective Recomputation for Efficient KV Cache Reuse in Retrieval-Augmented GenerationShihao Wang, Jiahao Chen, Yanqi Pan, Hao Huang et al.ICML 2026 · 4 citations
- CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge FusionJiayi Yao, Hanchen Li, Yuhan Liu, Siddhant Ray et al.EuroSys 2025 · 68 citations
- AdaCache: Adaptive Caching and Context Augmentation for Efficient LLM ServingZihao Zeng, Siyi Li, Xinyu Yan, Lei Xiao et al.ICLR 2026
- Accelerating Inference of Retrieval-Augmented Generation via Sparse Context SelectionYun Zhu, Jia-Chen Gu, Caitlin Sikora, Ho Ko et al.ICLR 2025
