Neural Path Hunter: Reducing Hallucination in Dialogue Systems via Path Grounding
Nouha Dziri, Andrea Madotto, Osmar Zaïane, Avishek Joey Bose
Abstract
Dialogue systems powered by large pretrained language models exhibit an innate ability to deliver fluent and natural-sounding responses. Despite their impressive performance, these models are fitful and can often generate factually incorrect statements impeding their widespread adoption. In this paper, we focus on the task of improving faithfulness and reducing hallucination of neural dialogue systems to known facts supplied by a Knowledge Graph (KG). We propose NEU-RAL PATH HUNTER which follows a generatethen-refine strategy whereby a generated response is amended using the KG. NEURAL PATH HUNTER leverages a separate tokenlevel fact critic to identify plausible sources of hallucination followed by a refinement stage that retrieves correct entities by crafting a query signal that is propagated over a k-hop subgraph. We empirically validate our proposed approach on the OpenDialKG dataset (Moon et al., 2019) against a suite of metrics and report a relative improvement of faithfulness over dialogue responses by 20.35% based on FeQA (Durmus et al.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers28
- Factuality Enhanced Language Models for Open-Ended Text GenerationNayeon Lee, Wei Ping, Peng Xu, Mostofa Patwary et al.NeurIPS 2022 · 318 citations
- HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language ModelsJunyi Li, Xiaoxue Cheng, Xin Zhao, Jian-Yun Nie et al.EMNLP 2023 · 224 citations
- Mathemyths: Leveraging Large Language Models to Teach Mathematical Language through Child-AI Co-Creative StorytellingChao Zhang, Xuechen Liu, Katherine Ziska, Soobin Jeon et al.CHI 2024 · 93 citations
- Contrastive Learning Reduces Hallucination in ConversationsWeiwei Sun, Zhengliang Shi, Shen Gao, Pengjie Ren et al.AAAI 2023 · 92 citations
- Exposing Attention Glitches with Flip-Flop Language ModelingBingbin Liu, Jordan T. Ash, Surbhi Goel, Akshay Krishnamurthy et al.NeurIPS 2023 · 90 citations
Builds on8
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Composition-based Multi-Relational Graph Convolutional NetworksShikhar Vashishth, Soumya Sanyal, Vikram Nitin, Partha P. TalukdarICLR 2020 · 1,105 citations
- Asking and Answering Questions to Evaluate the Factual Consistency of SummariesAlex Wang, Kyunghyun Cho, Mike LewisACL 2020 · 317 citations
- FEQA: A Question Answering Evaluation Framework for Faithfulness Assessment in Abstractive SummarizationEsin Durmus, He He, Mona T. DiabACL 2020 · 90 citations
- ToTTo: A Controlled Table-To-Text Generation DatasetAnkur P. Parikh, Xuezhi Wang, Sebastian Gehrmann, Manaal Faruqui et al.EMNLP 2020 · 69 citations
Related papers
- Post-hoc Utterance Refining Method by Entity Mining for Faithful Knowledge Grounded ConversationsYoonna Jang, Suhyune Son, Jeongwoo Lee, Junyoung Son et al.EMNLP 2023
- Deliberation on Priors: Trustworthy Reasoning of Large Language Models on Knowledge GraphsJie Ma, Ning Qu, Zhitao Gao, Rui Xing et al.NeurIPS 2025 · 9 citations
- MultiHal: Multilingual Dataset for Knowledge-Graph Grounded Evaluation of LLM HallucinationsErnests Lavrinovics, Russa Biswas, Katja Hose, Johannes BjervaICML 2026 · 5 citations
- Mitigating Large Language Model Hallucinations via Autonomous Knowledge Graph-Based RetrofittingXinyan Guan, Yanjiang Liu, Hongyu Lin, Yaojie Lu et al.AAAI 2024 · 127 citations
- KnowGPT: Knowledge Graph based Prompting for Large Language ModelsQinggang Zhang, Junnan Dong, Hao Chen, Daochen Zha et al.NeurIPS 2024 · 66 citations
