Scaling Knowledge Graph Construction through Synthetic Data Generation and Distillation
Prafulla Kumar Choubey, Xin Su, Man Luo, XIANGYU PENG, Caiming Xiong, Tiep Le, Shachar Rosenman, Vasudev Lal, Phil L Mui, Ricky Ho, Phillip Howard, Chien-Sheng Wu
Abstract
Document-level knowledge graph (KG) construction faces a fundamental scaling challenge: existing methods either rely on expensive large language models (LLMs), making them economically nonviable for large-scale corpora, or employ smaller models that produce incomplete and inconsistent graphs. We find that this limitation stems not from model capabilities but from insufficient training on high-quality document-level KG data. To address this gap, we introduce SynthKG, a multi-step data synthesis pipeline that generates high-quality document-KG pairs through systematic chunking, decontextualization, and structured extraction using LLMs. By fine-tuning a smaller LLM on synthesized document-KG pairs, we streamline the multi-step process into a single-step KG generation approach called Distill-SynthKG. Furthermore, we repurpose existing question-answering datasets to construct KG evaluation datasets and introduce new evaluation metrics. Using KGs produced by Distill-SynthKG, we also design a novel graph-based retrieval framework for RAG. Experimental results demonstrate that Distill-SynthKG not only surpasses all baseline models in KG quality (including models up to eight times larger) but also consistently improves in retrieval and question-answering tasks. Additionally, our proposed graph retrieval framework outperforms all KG-retrieval methods across multiple benchmark datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on6
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- HippoRAG: Neurobiologically Inspired Long-Term Memory for Large Language ModelsBernal Jimenez Gutierrez, Yiheng Shu, Yu Gu, Michihiro Yasunaga et al.NeurIPS 2024 · 395 citations
- Consistency Guided Knowledge Retrieval and Denoising in LLMs for Zero-shot Document-level Relation Triplet ExtractionQi Sun, Kun Huang, Xiaocui Yang, Rong Tong et al.WWW 2024 · 40 citations
- Schema-aware Reference as Prompt Improves Data-Efficient Knowledge Graph ConstructionYunzhi Yao, Shengyu Mao, Ningyu Zhang, Xiang Chen et al.SIGIR 2023 · 23 citations
Related papers
- Stepwise Contrastive Reasoning for Retrieval-Augmented Generation over Knowledge GraphsChenxiao Lin, Ye Luo, Kunhong Liu, Qingqiang WuAAAI 2026
- KET-RAG: A Cost-Efficient Multi-Granular Indexing Framework for Graph-RAGYiqian Huang, Shiqi Zhang, Xiaokui XiaoKDD 2025 · 6 citations
- Large Language Models Meet Knowledge Graphs for Question Answering: Synthesis and OpportunitiesChuangtao Ma, Yongrui Chen, Tianxing Wu, Arijit Khan et al.EMNLP 2025 · 6 citations
- Enrich-on-Graph: Query-Graph Alignment for Complex Reasoning with LLM EnrichingSongze Li, Zhiqiang Liu, Zhengke Gui, Huajun Chen et al.EMNLP 2025 · 6 citations
- Simple is Effective: The Roles of Graphs and Large Language Models in Knowledge-Graph-Based Retrieval-Augmented GenerationMufei Li, Siqi Miao, Pan LiICLR 2025
