CypherSmith: Transforming Text-to-Cypher Generation for LLMs with Synthetic Data
Zeyu Zhang, Kexuan Sun, Zheng Tang, Jens-S. Vöckler, Thien Huu Nguyen, Thuy Vu
摘要
Knowledge Graph (KG) retrieval is a promising augmentation to address knowledge gaps and hallucinations in LLMs. As KGs in practice are stored in graph databases (e.g., Wikidata, Freebase), accurate retrieval requires translating natural language questions into structured queries (query generation). A key challenge of query generation is Text-to-Cypher, which generates Cypher queries for property graphs (e.g., Neo4j), a paradigm increasingly adopted in industry for their scalable architectures and expressive schemas. However, compared to other query generation tasks such as Text-to-SQL or Text-to-SPARQL, Text-to-Cypher remains underexplored due to scarce public KGs and datasets. Existing datasets are small, domain-limited, and lack diversity, constraining LLM progress. To address this, we introduce CypherSmith, an instruction-tuning dataset over 12× larger than prior public Textto-Cypher datasets, spanning diverse domains to better support LLM fine-tuning. Our key distinction lies in fully leveraging open-source LLMs for large-scale synthetic data generation and introducing a novel likelihood-based filtering technique to ensure high-quality Text-to-Cypher data. Extensive experiments demonstrate the effectiveness of CypherSmith, achieving state-of-the-art LLM performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper20
- LIMA: Less Is More for AlignmentChunting Zhou, Pengfei Liu, Puxin Xu, Srinivasan Iyer 等NeurIPS 2023 · 被引用 1,486 次
- Reasoning on Graphs: Faithful and Interpretable Large Language Model ReasoningLinhao Luo, Yuan-Fang Li, Gholamreza Haffari, Shirui PanICLR 2024 · 被引用 499 次
- Beyond I.I.D.: Three Levels of Generalization for Question Answering on Knowledge BasesYu Gu, Sue Kase, Michelle Vanni, Brian M. Sadler 等WWW 2021 · 被引用 304 次
- AlpaGasus: Training a Better Alpaca with Fewer DataLichang Chen, Shiyang Li, Jun Yan, Hai Wang 等ICLR 2024 · 被引用 295 次
- #InsTag: Instruction Tagging for Analyzing Supervised Fine-tuning of Large Language ModelsKeming Lu, Hongyi Yuan, Zheng Yuan, Runji Lin 等ICLR 2024 · 被引用 119 次
相关 Paper
- CypherBench: Towards Precise Retrieval over Full-scale Modern Knowledge Graphs in the LLM EraYanlin Feng, Simone Papicchio, Sajjadur RahmanACL 2025
- Cypher-RI: Reinforcement Learning for Integrating Schema Selection into Cypher GenerationHanchen Su, Xuyuan Li, Yan Zhou, Zhuoyi Lu 等NeurIPS 2025 · 被引用 1 次
- GDsmith: Detecting Bugs in Cypher Graph Database EnginesZiyue Hua, Wei Lin, Luyao Ren, Zongyang Li 等ISSTA 2023 · 被引用 24 次
- Large Language Models Meet Knowledge Graphs for Question Answering: Synthesis and OpportunitiesChuangtao Ma, Yongrui Chen, Tianxing Wu, Arijit Khan 等EMNLP 2025 · 被引用 6 次
- Scaling Knowledge Graph Construction through Synthetic Data Generation and DistillationPrafulla Kumar Choubey, Xin Su, Man Luo, XIANGYU PENG 等ICLR 2026 · 被引用 5 次
