Detecting and Reducing the Factual Hallucinations of Large Language Models with Metamorphic Testing
Weibin Wu, Yuhang Cao, Ning Yi, Rongyi Ou, Zibin Zheng
Abstract
Question answering (QA) is a fundamental task of large language models (LLMs), which requires LLMs to automatically answer human-posed questions in natural language. However, LLMs are known to distort facts and make non-factual statements (i.e., hallucinations) when dealing with QA tasks, which may affect the deployment of LLMs in real-life situations. In this work, we propose DrHall, a framework for detecting and reducing the factual hallucinations of black-box LLMs with metamorphic testing (MT). We believe that hallucinated answers are unstable. Therefore, when LLMs hallucinate, they are more likely to produce different answers if we use metamorphic testing to make LLMs re-execute the same task with different execution paths, which motivates the design of DrHall. The effectiveness of DrHall is evaluated empirically on three datasets, including a self-built dataset of natural language questions: FactHalluQA, as well as two datasets of programming questions: Refactory and LeetCode. The evaluation results confirm that DrHall can consistently outperform the state-of-the-art baselines, obtaining an average F1 score of over 0.856 for hallucination detection. For hallucination correction, DrHall can also outperform the state-of-the-art baselines, with an average hallucination correction rate of over 53%. We hope that our work can enhance the reliability of LLMs and provide new insights for the research of LLM hallucination mitigation.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 74beb6a5-dec0-413b-a86d-146dc729538eCited by top-tier papers3
- Ref4D-VideoBench: Four-Dimensional Reference-Based Evaluation of Text-to-Video Generative ModelsJiajia Wei, YuJia He, Yuhan Hou, Hang Qi et al.CVPR 2026
- Validating LLM-Generated SQL Queries through Metamorphic PromptingLi Lin, Qinglin Zhu, Jintai Hong, Chong Wang et al.FSE 2026
- QRShield: Exploiting Vulnerabilities of Latent Diffusion Models for Preventing AI Art PlagiarismXunyue Mo, Weibin Wu, Qingrui Tu, Hang Wang et al.AAAI 2026
Related papers
- Hallucination Detection in Large Language Models with Metamorphic RelationsBorui Yang, Md Afif Al Mamun, Jie M. Zhang, Gias UddinFSE 2025 · 14 citations
- FactTest: Factuality Testing in Large Language Models with Finite-Sample and Distribution-Free GuaranteesFan Nie, Xiaotian Hou, Shuhang Lin, James Zou et al.ICML 2025
- Drowzee: Metamorphic Testing for Fact-Conflicting Hallucination Detection in Large Language ModelsNingke Li, Yuekang Li, Yi Liu, Ling Shi et al.OOPSLA 2024 · 26 citations
- HalluClean: A Unified Framework to Combat Hallucinations in LLMsYaxin Zhao, Yu ZhangAAAI 2026
- The Dawn After the Dark: An Empirical Study on Factuality Hallucination in Large Language ModelsJunyi Li, Jie Chen, Ruiyang Ren, Xiaoxue Cheng et al.ACL 2024 · 49 citations
