Scaling Knowledge Graph Construction through Synthetic Data Generation and Distillation
Prafulla Kumar Choubey, Xin Su, Man Luo, XIANGYU PENG, Caiming Xiong, Tiep Le, Shachar Rosenman, Vasudev Lal, Phil L Mui, Ricky Ho, Phillip Howard, Chien-Sheng Wu
摘要
Document-level knowledge graph (KG) construction faces a fundamental scaling challenge: existing methods either rely on expensive large language models (LLMs), making them economically nonviable for large-scale corpora, or employ smaller models that produce incomplete and inconsistent graphs. We find that this limitation stems not from model capabilities but from insufficient training on high-quality document-level KG data. To address this gap, we introduce SynthKG, a multi-step data synthesis pipeline that generates high-quality document-KG pairs through systematic chunking, decontextualization, and structured extraction using LLMs. By fine-tuning a smaller LLM on synthesized document-KG pairs, we streamline the multi-step process into a single-step KG generation approach called Distill-SynthKG. Furthermore, we repurpose existing question-answering datasets to construct KG evaluation datasets and introduce new evaluation metrics. Using KGs produced by Distill-SynthKG, we also design a novel graph-based retrieval framework for RAG. Experimental results demonstrate that Distill-SynthKG not only surpasses all baseline models in KG quality (including models up to eight times larger) but also consistently improves in retrieval and question-answering tasks. Additionally, our proposed graph retrieval framework outperforms all KG-retrieval methods across multiple benchmark datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- HippoRAG: Neurobiologically Inspired Long-Term Memory for Large Language ModelsBernal Jimenez Gutierrez, Yiheng Shu, Yu Gu, Michihiro Yasunaga 等NeurIPS 2024 · 被引用 395 次
- Consistency Guided Knowledge Retrieval and Denoising in LLMs for Zero-shot Document-level Relation Triplet ExtractionQi Sun, Kun Huang, Xiaocui Yang, Rong Tong 等WWW 2024 · 被引用 40 次
- Schema-aware Reference as Prompt Improves Data-Efficient Knowledge Graph ConstructionYunzhi Yao, Shengyu Mao, Ningyu Zhang, Xiang Chen 等SIGIR 2023 · 被引用 23 次
相关 Paper
- Stepwise Contrastive Reasoning for Retrieval-Augmented Generation over Knowledge GraphsChenxiao Lin, Ye Luo, Kunhong Liu, Qingqiang WuAAAI 2026
- KET-RAG: A Cost-Efficient Multi-Granular Indexing Framework for Graph-RAGYiqian Huang, Shiqi Zhang, Xiaokui XiaoKDD 2025 · 被引用 6 次
- Large Language Models Meet Knowledge Graphs for Question Answering: Synthesis and OpportunitiesChuangtao Ma, Yongrui Chen, Tianxing Wu, Arijit Khan 等EMNLP 2025 · 被引用 6 次
- Enrich-on-Graph: Query-Graph Alignment for Complex Reasoning with LLM EnrichingSongze Li, Zhiqiang Liu, Zhengke Gui, Huajun Chen 等EMNLP 2025 · 被引用 6 次
- Simple is Effective: The Roles of Graphs and Large Language Models in Knowledge-Graph-Based Retrieval-Augmented GenerationMufei Li, Siqi Miao, Pan LiICLR 2025
