SubGCache: Accelerating Graph-based RAG with Subgraph-level KV Cache
Qiuyu Zhu, Liang Zhang, Qianxiong Xu, Cheng Long, Jie Zhang
Abstract
Graph-based retrieval-augmented generation (RAG) enables large language models (LLMs) to incorporate structured knowledge via graph retrieval as contextual input, enhancing more accurate and context-aware reasoning. We observe that for different queries, it could retrieve similar subgraphs as prompts, and thus we propose SubGCache, which aims to reduce inference latency by reusing computation across queries with similar structural prompts (i.e., subgraphs). Specifically, SubGCache clusters queries based on subgraph embeddings, constructs a representative subgraph for each cluster, and pre-computes the key-value (KV) cache of the representative subgraph. For each query with its retrieved subgraph within a cluster, it reuses the pre-computed KV cache of the representative subgraph of the cluster without computing the KV tensors again for saving computation. Extensive experiments on three datasets across multiple LLM backbones and graph-based RAG frameworks demonstrate that SubGCache consistently reduces inference latency with comparable and even improved generation quality, achieving up to 6.68x reduction in time-to-first-token (TTFT).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cd192046-b4df-4a73-8c87-0dc9b4b42f7bBuilds on10
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu et al.NeurIPS 2023 · 5,989 citations
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionAkari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil et al.ICLR 2024 · 1,798 citations
- Improving Language Models by Retrieving from Trillions of TokensSebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai et al.ICML 2022 · 1,629 citations
- Can Language Models Solve Graph Problems in Natural Language?Heng Wang, Shangbin Feng, Tianxing He, Zhaoxuan Tan et al.NeurIPS 2023 · 420 citations
Related papers
- DepCache: A KV Cache Management Framework for GraphRAG with Dependency AttentionHao Yuan, Xin Ai, Qiange Wang, Peizheng Li et al.SIGMOD 2026 · 2 citations
- Structure Is All You Need to Reuse: Accelerating GraphRAG via Meta-Structure-Aware KV CachingRuikun Luo, Changwei Gu, Jing Yang, Hongming Liang et al.KDD 2026
- TurboRAG: Accelerating Retrieval-Augmented Generation with Precomputed KV Caches for Chunked TextSongshuo Lu, Hua Wang, Yutian Rong, Zhi Chen et al.EMNLP 2025 · 2 citations
- Simple is Effective: The Roles of Graphs and Large Language Models in Knowledge-Graph-Based Retrieval-Augmented GenerationMufei Li, Siqi Miao, Pan LiICLR 2025
- Cache-Craft: Managing Chunk-Caches for Efficient Retrieval-Augmented GenerationShubham Agarwal, Sai Sundaresan, Subrata Mitra, Debabrata Mahapatra et al.SIGMOD 2025 · 20 citations
