Cache-Craft: Managing Chunk-Caches for Efficient Retrieval-Augmented Generation
Shubham Agarwal, Sai Sundaresan, Subrata Mitra, Debabrata Mahapatra, Archit Gupta, Rounak Sharma, Nirmal Joshua Kapu, Tong Yu, Shiv Kumar Saini
摘要
Retrieval-Augmented Generation (RAG) is often used with Large Language Models (LLMs) to infuse domain knowledge or user-specific information. In RAG, given a user query, a retriever extracts chunks of relevant text from a knowledge base. These chunks are sent to an LLM as part of the input prompt. Typically, any given chunk is repeatedly retrieved across user questions. However, currently, for every question, attention layers in LLMs fully compute the Keys and Values (KVs) repeatedly for the input chunks, as state-of-the-art methods cannot reuse KV-caches when chunks appear at arbitrary locations or with arbitrary contexts. Naive reuse leads to output quality degradation. This leads to potentially redundant computations on expensive GPUs and increases latency. In this work, we propose Cache-Craft , a system for managing and reusing precomputed KVs corresponding to the text chunks (which we call chunk-caches ) in RAG-based systems. We present how to identify chunk-caches that are reusable, how to efficiently perform a small fraction of recomputation to fix the cache and maintain output quality, and how to efficiently store and evict chunk-caches in the hardware for maximizing reuse while masking any overheads. With real production workloads as well as synthetic datasets, we show that Cache-Craft reduces redundant computation by 51% over SOTA prefix-caching and 75% over full recomputation. Additionally, with continuous batching on a real production workload, we get a 1.6× speed up in throughput for both the LLama-3-8B and 70B models and a 2.1× and 2× reduction in end-to-end response latency respectively, compared to prefix-caching, while maintaining generation quality.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- In-depth Analysis of Graph-based RAG in a Unified FrameworkYingli Zhou, Yaodong Su, Youran Sun, Shu Wang 等VLDB 2025 · 被引用 48 次
- DualMap: Enabling Both Cache Affinity and Load Balancing for Distributed LLM ServingYing Yuan, Pengfei Zuo, Bo Wang, Zhangyu Chen 等ICLR 2026 · 被引用 10 次
- CoDec: Prefix-Shared Decoding Kernel for LLMsZhibin Wang, Rui Ning, Chao Fang, Zhonghui Zhang 等SIGMOD 2026 · 被引用 8 次
- Generative Caching for Structurally Similar Prompts and ResponsesSarthak Chakraborty, Suman Nath, Xuchao Zhang, Chetan Bansal 等NeurIPS 2025 · 被引用 5 次
- MIRAGE: Misleading Retrieval-Augmented Generation via Black-box and Query-agnostic Poisoning AttacksTailun Chen, Yu He, Yan Wang, Shuo Shao 等CCS 2026 · 被引用 4 次
它引用的顶会 Paper33
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- SparseGPT: Massive Language Models Can be Accurately Pruned in One-ShotElias Frantar, Dan AlistarhICML 2023 · 被引用 1,240 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language ModelsZhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen 等NeurIPS 2023 · 被引用 1,003 次
- KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache QuantizationColeman Hooper, Sehoon Kim, Hiva Mohammadzadeh, Michael W. Mahoney 等NeurIPS 2024 · 被引用 738 次
相关 Paper
- CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge FusionJiayi Yao, Hanchen Li, Yuhan Liu, Siddhant Ray 等EuroSys 2025 · 被引用 68 次
- From Prefix Cache to Fusion RAG Cache: Accelerating LLM Inference in Retrieval-Augmented GenerationJiahao Wang, Weiyu Xie, Mingxing Zhang, Boxin Zhang 等SIGMOD 2026 · 被引用 4 次
- SubGCache: Accelerating Graph-based RAG with Subgraph-level KV CacheQiuyu Zhu, Liang Zhang, Qianxiong Xu, Cheng Long 等AAAI 2026 · 被引用 1 次
- AdaCache: Adaptive Caching and Context Augmentation for Efficient LLM ServingZihao Zeng, Siyi Li, Xinyu Yan, Lei Xiao 等ICLR 2026
- TurboRAG: Accelerating Retrieval-Augmented Generation with Precomputed KV Caches for Chunked TextSongshuo Lu, Hua Wang, Yutian Rong, Zhi Chen 等EMNLP 2025 · 被引用 2 次
