MultiHal: Multilingual Dataset for Knowledge-Graph Grounded Evaluation of LLM Hallucinations
Ernests Lavrinovics, Russa Biswas, Katja Hose, Johannes Bjerva
摘要
Large Language Models (LLMs) have inherent limitations of faithfulness and factuality, commonly referred to as hallucinations. Several benchmarks have been developed that provide a test bed for factuality evaluation within the context of English-centric datasets, while relying on supplementary informative context like web links or text passages but ignoring the available structured factual resources. To this end, Knowledge Graphs (KGs) have been identified as a useful aid for hallucination mitigation, as they provide a structured way to represent the facts about entities and their relations with minimal linguistic overhead. We bridge the lack of KG paths and multilinguality for factual language modeling within the existing hallucination evaluation benchmarks and propose a KG-based multilingual, multihop benchmark called MultiHal framed for generative text evaluation. As part of our data collection pipeline, we mined 140k KG-paths from open-domain KGs, from which we pruned noisy KG-paths, curating a high-quality subset of 25.9k. Our baseline evaluation shows an absolute scale improvement by approximately 0.12 to 0.36 points for the semantic similarity score, 0.16 to 0.36 for NLI entailment and 0.29 to 0.42 for hallucination detection in KG-RAG over vanilla QA across multiple languages and multiple models, demonstrating the potential of KG integration. We anticipate MultiHal will foster future research towards several graph-based hallucination mitigation and fact-checking tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- Reasoning on Graphs: Faithful and Interpretable Large Language Model ReasoningLinhao Luo, Yuan-Fang Li, Gholamreza Haffari, Shirui PanICLR 2024 · 被引用 499 次
- RankRAG: Unifying Context Ranking with Retrieval-Augmented Generation in LLMsYue Yu, Wei Ping, Zihan Liu, Boxin Wang 等NeurIPS 2024 · 被引用 321 次
- Think-on-Graph: Deep and Responsible Reasoning of Large Language Model on Knowledge GraphJiashuo Sun, Chengjin Xu, Lumingyuan Tang, Saizhuo Wang 等ICLR 2024 · 被引用 247 次
- HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language ModelsJunyi Li, Xiaoxue Cheng, Xin Zhao, Jian-Yun Nie 等EMNLP 2023 · 被引用 224 次
- Mitigating Large Language Model Hallucinations via Autonomous Knowledge Graph-Based RetrofittingXinyan Guan, Yanjiang Liu, Hongyu Lin, Yaojie Lu 等AAAI 2024 · 被引用 127 次
相关 Paper
- Can Knowledge Graphs Make Large Language Models More Trustworthy? An Empirical Study Over Open-ended Question AnsweringYuan Sui, Yufei He, Zifeng Ding, Bryan HooiACL 2025 · 被引用 29 次
- K-HALU: Multiple Answer Korean Hallucination Benchmark for Large Language ModelsJaehyung Seo, Heuiseok LimICLR 2025
- KnowGPT: Knowledge Graph based Prompting for Large Language ModelsQinggang Zhang, Junnan Dong, Hao Chen, Daochen Zha 等NeurIPS 2024 · 被引用 66 次
- QA-GraphRAG: Query-Adaptive Plug-and-Play Retrieval Integration for Graph-based Retrieval-Augmented GenerationZeang Sheng, Ruihong Sun, Jiahao Xu, Hanmei Luo 等VLDB 2026
- GraphCheck: Breaking Long-Term Text Barriers with Extracted Knowledge Graph-Powered Fact-CheckingYingjian Chen, Haoran Liu, Yinhong Liu, Jinxiang Xie 等ACL 2025
