Hallucination Mitigation in Natural Language Generation from Large-Scale Open-Domain Knowledge Graphs
Xiao Shi, Zhengyuan Zhu, Zeyu Zhang, Chengkai Li
Abstract
In generating natural language descriptions for knowledge graph triples, prior works used either small-scale, human-annotated datasets or datasets with limited variety of graph shapes, e.g., those having mostly star graphs. Graph-to-text models trained and evaluated on such datasets are largely not assessed for more realistic large-scale, open-domain settings. We introduce a new dataset, GraphNarrative, to fill this gap. Fine-tuning transformer-based pre-trained language models has achieved state-of-the-art performance among graph-to-text models. However, this method suffers from information hallucination—the generated text may contain fabricated facts not present in input graphs. We propose a novel approach that, given a graph-sentence pair in GraphNarrative, trims the sentence to eliminate portions that are not present in the corresponding graph, by utilizing the sentence’s dependency parse tree. Our experiment results verify this approach using models trained on GraphNarrative and existing datasets. The dataset, source code, and trained models are released at https://github.com/idirlab/graphnarrator.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Debate on Graph: A Flexible and Reliable Reasoning Framework for Large Language ModelsJie Ma, Zhitao Gao, Qi Chai, Wangchun Sun et al.AAAI 2025 · 8 citations
- An Audit on the Perspectives and Challenges of Hallucinations in NLPPranav Narayanan Venkit, Tatiana Chakravorti, Vipul Gupta, Heidi Biggs et al.EMNLP 2024 · 8 citations
Builds on4
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- HTLM: Hyper-Text Pre-Training and Prompting of Language ModelsArmen Aghajanyan, Dmytro Okhonko, Mike Lewis, Mandar Joshi et al.ICLR 2022 · 82 citations
- Open Domain Question Answering with A Unified Knowledge InterfaceKaixin Ma, Hao Cheng, Xiaodong Liu, Eric Nyberg et al.ACL 2022 · 45 citations
Related papers
- Structure-aware Knowledge Graph-to-text Generation with Planning Selection and Similarity DistinctionFeng Zhao, Hongzhi Zou, Cheng YanEMNLP 2023 · 5 citations
- ScriptWriter: Narrative-Guided Script GenerationYutao Zhu, Ruihua Song, Zhicheng Dou, Jian-Yun Nie et al.ACL 2020 · 23 citations
- DiscoSG: Towards Discourse-Level Text Scene Graph Parsing through Iterative Graph RefinementShaoqing Lin, Chong Teng, Fei Li, Donghong Ji et al.EMNLP 2025
- Generating Coherent Narratives by Learning Dynamic and Discrete Entity States with a Contrastive FrameworkJian Guan, Zhenyu Yang, Rongsheng Zhang, Zhipeng Hu et al.AAAI 2023 · 11 citations
- ENT-DESC: Entity Description Generation by Exploring Knowledge GraphLiying Cheng, Dekun Wu, Lidong Bing, Yan Zhang et al.EMNLP 2020 · 22 citations
