Graph-KV: Breaking Sequence via Injecting Structural Biases into Large Language Models
Haoyu Wang, Peihao Wang, Mufei Li, Shikun Liu, Siqi Miao, Zhangyang (Atlas) Wang, Pan Li
Abstract
Modern large language models (LLMs) are inherently auto-regressive, requiring input to be serialized into flat sequences regardless of their structural dependencies. This serialization hinders the model's ability to leverage structural inductive biases, especially in tasks such as retrieval-augmented generation (RAG) and reasoning on data with native graph structures, where inter-segment dependencies are crucial. We introduce Graph-KV with the potential to overcome this limitation. Graph-KV leverages the KV-cache of text segments as condensed representations and governs their interaction through structural inductive biases. In this framework, "target" segments selectively attend only to the KV-caches of their designated "source" segments, rather than all preceding segments in a serialized sequence. This approach induces a graph-structured block mask, sparsifying attention and enabling a message-passing-like step within the LLM. Furthermore, strategically allocated positional encodings for source and target segments reduce positional bias and context window consumption. We evaluate Graph-KV across three scenarios: (1) seven RAG benchmarks spanning direct inference, multi-hop reasoning, and long-document understanding; (2) ARXIV-QA, a novel academic paper QA task with full-text scientific papers structured as citation ego-graphs; and (3) paper topic classification within a citation network. By effectively reducing positional bias and harnessing structural inductive biases, Graph-KV substantially outperforms baselines, including standard costly sequential encoding, across various settings. Code and the ARXIV-QA data are publicly available at https://github.com/ Graph-COM/GraphKV.
To address these limitations, we introduce Graph-KV. The core principle of Graph-KV is to treat the KV cache of a given text segment as its condensed information representation and to control its generation using structural inductive biases. Specifically, after initially prefilling the independent KV caches of all text segments, a "target" segment's KV cache is generated by attending only to the KV caches of its "source" text segments, rather than to all segments that merely precede it in the token-serialization sequence. The determination of "source → target" relationships is guided by structural inductive biases tied to either the data or the specific tasks. From another perspective, this approach essentially introduces a graph-structured block mask (Fig. 1) that sparsifies attention computation during KV cache generation, effectively enabling a "message passing through graph" step within the LLM. Moreover, to mitigate inherent positional biases from the LLM, the attention computation imposes shared PEs across the source segments, with the target segment receiving PEs with position indices immediately following its sources. This design substantially reduces context window consumption through shared PEs while preserving structural alignment.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a7d66ffc-fb99-4982-8438-8d6595eec24fCited by top-tier papers2
- Weaving Graph over Tokens: Contextualizing Structured Sequences for LLMsJiaxuan Chen, Zixing Zhang, Ruijun Mao, Wei Sun et al.ICML 2026
- SLASH the Sink: Sharpening Structural Attention Inside LLMsYiming Liu, Bin Lu, Xinbing Wang, Chenghu Zhou et al.ICML 2026
Builds on27
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong et al.NeurIPS 2020 · 3,935 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Calibrate Before Use: Improving Few-shot Performance of Language ModelsZihao Zhao, Eric Wallace, Shi Feng, Dan Klein et al.ICML 2021 · 1,843 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
Related papers
- Structure Is All You Need to Reuse: Accelerating GraphRAG via Meta-Structure-Aware KV CachingRuikun Luo, Changwei Gu, Jing Yang, Hongming Liang et al.KDD 2026
- DepCache: A KV Cache Management Framework for GraphRAG with Dependency AttentionHao Yuan, Xin Ai, Qiange Wang, Peizheng Li et al.SIGMOD 2026 · 2 citations
- LazyAttention: Efficient Retrieval-Augmented Generation with Deferred Positional EncodingHaocheng Xia, Mihir Pamnani, Hanxi Fang, Supawit Chockchowwat et al.ICML 2026
- SubGCache: Accelerating Graph-based RAG with Subgraph-level KV CacheQiuyu Zhu, Liang Zhang, Qianxiong Xu, Cheng Long et al.AAAI 2026 · 1 citation
- GraphKV: Breaking the Static Selection Paradigm with Graph-Based KV Cache EvictionXuelin Li, Xiangqi Jin, Linfeng ZhangEMNLP 2025
