CypherBench: Towards Precise Retrieval over Full-scale Modern Knowledge Graphs in the LLM Era
Yanlin Feng, Simone Papicchio, Sajjadur Rahman
Abstract
Retrieval from graph data is crucial for augmenting large language models (LLM) with both open-domain knowledge and private enterprise data, and it is also a key component in the recent GraphRAG system [1] . Despite decades of research on knowledge graphs and knowledge base question answering, leading LLM frameworks (e.g., Langchain and LlamaIndex) have only minimal support for retrieval from modern encyclopedic knowledge graphs like Wikidata. In this paper, we analyze the root cause and suggest that modern RDF knowledge graphs (e.g., Wikidata, Freebase) are less efficient for LLMs due to overly large schemas that far exceed the typical LLM context window, use of resource identifiers, overlapping relation types and lack of normalization. As a solution, we propose property graph views on top of the underlying RDF graph that can be efficiently queried by LLMs using Cypher. We instantiated this idea on Wikidata and introduced CypherBench, the first benchmark with 11 large-scale, multi-domain property graphs with 7.8 million entities and over 10,000 questions. To achieve this, we tackled several key challenges, including developing an RDF-to-property graph conversion engine, creating a systematic pipeline for text-to-Cypher task generation, and designing new evaluation metrics. Dataset https://huggingface.co/datasets/megagonlabs/cypherbench Code https://github.com/megagonlabs/cypherbench * The work began during Simone Papicchio's internship at Megagon Labs. As part of one subtask of his overall internship goal, he implemented an initial version of the benchmark that involved SQL-inspired template design, query categorization, and validation of the generated benchmark. The work has since further evolved to broaden and bolster the template generation process and redefining query categories while introducing new evaluation metrics. 2 Graph retrieval can be considered as a broader task than KBQA, as it is not only essential for question answering but also for other tasks such as fact checking [9] .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1b590fc0-23a7-4894-a13f-4c92dadc7975Cited by top-tier papers5
- The Great Nugget Recall: Automating Fact Extraction and RAG Evaluation with Large Language ModelsRonak Pradeep, Nandan Thakur, Shivani Upadhyay, Daniel Campos et al.SIGIR 2025 · 13 citations
- Cypher-RI: Reinforcement Learning for Integrating Schema Selection into Cypher GenerationHanchen Su, Xuyuan Li, Yan Zhou, Zhuoyi Lu et al.NeurIPS 2025 · 1 citation
- GQLBench: A Large-Scale Cross-Domain, Cross-Dialect Benchmark for NL2GQLYanning Su, Yuhang Zhou, Yang Fang, Sen Liu et al.ACL 2026
- CypherSmith: Transforming Text-to-Cypher Generation for LLMs with Synthetic DataZeyu Zhang, Kexuan Sun, Zheng Tang, Jens-S. Vöckler et al.ACL 2026
- SQUAB: Evaluating LLM robustness to Ambiguous and Unanswerable Questions in Semantic ParsingSimone Papicchio, Luca Cagliero, Paolo PapottiEMNLP 2025
Builds on16
- Reasoning on Graphs: Faithful and Interpretable Large Language Model ReasoningLinhao Luo, Yuan-Fang Li, Gholamreza Haffari, Shirui PanICLR 2024 · 499 citations
- Beyond I.I.D.: Three Levels of Generalization for Question Answering on Knowledge BasesYu Gu, Sue Kase, Michelle Vanni, Brian M. Sadler et al.WWW 2021 · 304 citations
- Subgraph Retrieval Enhanced Model for Multi-hop Knowledge Base Question AnsweringJing Zhang, Xiaokang Zhang, Jifan Yu, Jian Tang et al.ACL 2022 · 221 citations
- Scalable Multi-Hop Relational Reasoning for Knowledge-Aware Question AnsweringYanlin Feng, Xinyue Chen, Bill Yuchen Lin, Peifeng Wang et al.EMNLP 2020 · 207 citations
- Paths-over-Graph: Knowledge Graph Empowered Large Language Model ReasoningXingyu Tan, Xiaoyang Wang, Qing Liu, Xiwei Xu et al.WWW 2025 · 86 citations
Related papers
- Evaluating LLMs on Large-Scale Graph Property Estimation via Random WalksSunil Kumar Maurya, Xin LiuACL 2026
- Dialogue Benchmark Generation from Knowledge Graphs with Cost-Effective Retrieval-Augmented LLMsReham Omar, Omij Mangukiya, Essam MansourSIGMOD 2025 · 9 citations
- LightPROF: A Lightweight Reasoning Framework for Large Language Model on Knowledge GraphTu Ao, Yanhua Yu, Yuling Wang, Yang Deng et al.AAAI 2025 · 28 citations
- When to use Graphs in RAG: A Comprehensive Analysis for Graph Retrieval-Augmented GenerationZhishang Xiang, Chuanjie Wu, Qinggang Zhang, Shengyuan Chen et al.ICLR 2026 · 56 citations
- CBench: Towards Better Evaluation of Question Answering Over Knowledge GraphsAbdelghny Orogat, Isabelle Liu, Ahmed El-RobyVLDB 2021 · 17 citations
