Structure Is All You Need to Reuse: Accelerating GraphRAG via Meta-Structure-Aware KV Caching
Ruikun Luo, Changwei Gu, Jing Yang, Hongming Liang, Yuan Gao, Xiaofen Wang, Qiang He, Song Wu, Laurence T. Yang, Hai Jin
Abstract
Retrieval-Augmented Generation over Knowledge Graphs (GraphRAG) enhances Large Language Models (LLMs) with structured, multi-hop evidence. However, existing GraphRAG systems predominantly linearize retrieved subgraphs into long textual prompts, forcing LLMs to recompute identical schema-level reasoning across queries repeatedly. This text-centric design incurs substantial prefilling latency, memory overhead, and severely limited cache reuse under entity-level variations. We observe that although retrieved entities differ across queries, their underlying logical schemas (meta-structures) recur with high frequency, indicating that most computational cost is spent on repeatedly encoding invariant structural logic. In this paper, we propose MetaKV, the first structure-aware KV caching mechanism that explicitly decouples static structural logic from dynamic entity semantics in GraphRAG inference. In a preparation phase, MetaKV mines frequent meta-structures and pre-computes their Key-Value (KV) caches as reusable Skeleton KVs. During inference, query-specific entity representations are injected into reserved structural slots to assemble the context without recomputing graph topology. To further enforce faithfulness to graph reasoning, MetaKV introduces a Topological Mask that constrains attention to valid graph edges. Extensive experiments conducted on HotpotQA and MetaQA datasets demonstrate that MetaKV achieves up to 6.4× prefilling speedup and a 73% effective cache-hit rate while maintaining competitive reasoning accuracy, enabling high-throughput, low-latency GraphRAG without sacrificing adherence to graph topology.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get abf34c9c-8dfd-425e-a149-5b1cf6754f1fRelated papers
- SubGCache: Accelerating Graph-based RAG with Subgraph-level KV CacheQiuyu Zhu, Liang Zhang, Qianxiong Xu, Cheng Long et al.AAAI 2026 · 1 citation
- DepCache: A KV Cache Management Framework for GraphRAG with Dependency AttentionHao Yuan, Xin Ai, Qiange Wang, Peizheng Li et al.SIGMOD 2026 · 2 citations
- You Don't Need Pre-Built Graphs for RAG: Retrieval Augmented Generation with Adaptive Reasoning StructuresShengyuan Chen, Chuang Zhou, Zheng Yuan, Qinggang Zhang et al.AAAI 2026 · 14 citations
- Graph-KV: Breaking Sequence via Injecting Structural Biases into Large Language ModelsHaoyu Wang, Peihao Wang, Mufei Li, Shikun Liu et al.NeurIPS 2025 · 6 citations
- MGRAG: Semantic Subgraph Matching and Graph-Aware Caching for Multimodal Retrieval-Augmented GenerationYubo Wang, Haoyang Li, Lei ChenVLDB 2026
