Drowzee: Metamorphic Testing for Fact-Conflicting Hallucination Detection in Large Language Models
Ningke Li, Yuekang Li, Yi Liu, Ling Shi, Kailong Wang, Haoyu Wang
Abstract
Large language models (LLMs) have revolutionized language processing, but face critical challenges with security, privacy, and generating hallucinations -coherent but factually inaccurate outputs. A major issue is fact-conflicting hallucination (FCH), where LLMs produce content contradicting ground truth facts. Addressing FCH is difficult due to two key challenges: 1) Automatically constructing and updating benchmark datasets is hard, as existing methods rely on manually curated static benchmarks that cannot cover the broad, evolving spectrum of FCH cases. 2) Validating the reasoning behind LLM outputs is inherently difficult, especially for complex logical relations.
To tackle these challenges, we introduce a novel logic-programming-aided metamorphic testing technique for FCH detection. We develop an extensive and extensible framework that constructs a comprehensive factual knowledge base by crawling sources like Wikipedia, seamlessly integrated into Drowzee 1 . Using logical reasoning rules, we transform and augment this knowledge into a large set of test cases with ground truth answers. We test LLMs on these cases through template-based prompts, requiring them to provide reasoned answers. To validate their reasoning, we propose two semantic-aware oracles that assess the similarity between the semantic structures of the LLM answers and ground truth. Our approach automatically generates useful test cases and identifies hallucinations across six LLMs within nine domains, with hallucination rates ranging from 24.7% to 59.8%. Key findings include LLMs struggling with temporal concepts, out-of-distribution knowledge, and lack of logical reasoning capabilities. The results show that logic-based test cases generated by Drowzee effectively trigger and detect hallucinations. To further mitigate the identified FCHs, we explored model editing techniques, which proved effective on a small scale (with edits to fewer than 1000 knowledge pieces). Our findings emphasize the need for continued community efforts to detect and mitigate model hallucinations.
CCS Concepts: • Software and its engineering → Software testing and debugging.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c67229e8-23d9-4983-91ad-065a9fbda6f9Cited by top-tier papers12
- RvLLM: LLM Runtime Verification with Domain KnowledgeYedi Zhang, Sun Yi Emma, Annabelle Lee Jia En, Jin Song DongNeurIPS 2025 · 19 citations
- Semantic-Enhanced Indirect Call Analysis with Large Language ModelsBaijun Cheng, Cen Zhang, Kailong Wang, Ling Shi et al.ASE 2024 · 4 citations
- Understanding the Effectiveness of Coverage Criteria for Large Language Models: A Special Angle from Jailbreak AttacksShide Zhou, Tianlin Li, Kailong Wang, Yihao Huang et al.ICSE 2025 · 3 citations
- GlitchProber: Advancing Effective Detection and Mitigation of Glitch Tokens in Large Language ModelsZhibo Zhang, Wuxia Bai, Yuxi Li, Mark Huasong Meng et al.ASE 2024 · 2 citations
- HFuzzer: Testing Large Language Models for Package Hallucinations via Phrase-based FuzzingYukai Zhao, Menghan Wu, Xing Hu, Xin XiaASE 2025 · 1 citation
Builds on18
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister et al.NeurIPS 2023 · 1,549 citations
- CRITIC: Large Language Models Can Self-Correct with Tool-Interactive CritiquingZhibin Gou, Zhihong Shao, Yeyun Gong, Yelong Shen et al.ICLR 2024 · 699 citations
Related papers
- Detecting and Reducing the Factual Hallucinations of Large Language Models with Metamorphic TestingWeibin Wu, Yuhang Cao, Ning Yi, Rongyi Ou et al.FSE 2025 · 5 citations
- Hallucination Detection in Large Language Models with Metamorphic RelationsBorui Yang, Md Afif Al Mamun, Jie M. Zhang, Gias UddinFSE 2025 · 14 citations
- InterrogateLLM: Zero-Resource Hallucination Detection in LLM-Generated AnswersYakir Yehuda, Itzik Malkiel, Oren Barkan, Jonathan Weill et al.ACL 2024
- The Dawn After the Dark: An Empirical Study on Factuality Hallucination in Large Language ModelsJunyi Li, Jie Chen, Ruiyang Ren, Xiaoxue Cheng et al.ACL 2024 · 49 citations
- Logical Consistency of Large Language Models in Fact-CheckingBishwamittra Ghosh, Sarah Hasan, Naheed Anjum Arafat, Arijit KhanICLR 2025
