CypherSmith: Transforming Text-to-Cypher Generation for LLMs with Synthetic Data
Zeyu Zhang, Kexuan Sun, Zheng Tang, Jens-S. Vöckler, Thien Huu Nguyen, Thuy Vu
Abstract
Knowledge Graph (KG) retrieval is a promising augmentation to address knowledge gaps and hallucinations in LLMs. As KGs in practice are stored in graph databases (e.g., Wikidata, Freebase), accurate retrieval requires translating natural language questions into structured queries (query generation). A key challenge of query generation is Text-to-Cypher, which generates Cypher queries for property graphs (e.g., Neo4j), a paradigm increasingly adopted in industry for their scalable architectures and expressive schemas. However, compared to other query generation tasks such as Text-to-SQL or Text-to-SPARQL, Text-to-Cypher remains underexplored due to scarce public KGs and datasets. Existing datasets are small, domain-limited, and lack diversity, constraining LLM progress. To address this, we introduce CypherSmith, an instruction-tuning dataset over 12× larger than prior public Textto-Cypher datasets, spanning diverse domains to better support LLM fine-tuning. Our key distinction lies in fully leveraging open-source LLMs for large-scale synthetic data generation and introducing a novel likelihood-based filtering technique to ensure high-quality Text-to-Cypher data. Extensive experiments demonstrate the effectiveness of CypherSmith, achieving state-of-the-art LLM performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e8b5c088-3031-4a50-98ce-e1c1e6cde16aBuilds on20
- LIMA: Less Is More for AlignmentChunting Zhou, Pengfei Liu, Puxin Xu, Srinivasan Iyer et al.NeurIPS 2023 · 1,486 citations
- Reasoning on Graphs: Faithful and Interpretable Large Language Model ReasoningLinhao Luo, Yuan-Fang Li, Gholamreza Haffari, Shirui PanICLR 2024 · 499 citations
- Beyond I.I.D.: Three Levels of Generalization for Question Answering on Knowledge BasesYu Gu, Sue Kase, Michelle Vanni, Brian M. Sadler et al.WWW 2021 · 304 citations
- AlpaGasus: Training a Better Alpaca with Fewer DataLichang Chen, Shiyang Li, Jun Yan, Hai Wang et al.ICLR 2024 · 295 citations
- #InsTag: Instruction Tagging for Analyzing Supervised Fine-tuning of Large Language ModelsKeming Lu, Hongyi Yuan, Zheng Yuan, Runji Lin et al.ICLR 2024 · 119 citations
Related papers
- CypherBench: Towards Precise Retrieval over Full-scale Modern Knowledge Graphs in the LLM EraYanlin Feng, Simone Papicchio, Sajjadur RahmanACL 2025
- Cypher-RI: Reinforcement Learning for Integrating Schema Selection into Cypher GenerationHanchen Su, Xuyuan Li, Yan Zhou, Zhuoyi Lu et al.NeurIPS 2025 · 1 citation
- GDsmith: Detecting Bugs in Cypher Graph Database EnginesZiyue Hua, Wei Lin, Luyao Ren, Zongyang Li et al.ISSTA 2023 · 24 citations
- Large Language Models Meet Knowledge Graphs for Question Answering: Synthesis and OpportunitiesChuangtao Ma, Yongrui Chen, Tianxing Wu, Arijit Khan et al.EMNLP 2025 · 6 citations
- Scaling Knowledge Graph Construction through Synthetic Data Generation and DistillationPrafulla Kumar Choubey, Xin Su, Man Luo, XIANGYU PENG et al.ICLR 2026 · 5 citations
