Hallucination Detection for Generative Large Language Models by Bayesian Sequential Estimation
Xiaohua Wang, Yuliang Yan, Longtao Huang, Xiaoqing Zheng, Xuanjing Huang
Abstract
Large Language Models (LLMs) have made remarkable advancements in the field of natural language generation. However, the propensity of LLMs to generate inaccurate or non-factual content, termed "hallucinations", remains a significant challenge. Current hallucination detection methods often necessitate the retrieval of great numbers of relevant evidence, thereby increasing response times. We introduce a unique framework that leverages statistical decision theory and Bayesian sequential analysis to optimize the trade-off between costs and benefits during the hallucination detection process. This approach does not require a predetermined number of observations. Instead, the analysis proceeds in a sequential manner, enabling an expeditious decision towards "belief" or "disbelief" through a stop-or-continue strategy. Extensive experiments reveal that this novel framework surpasses existing methods in both efficiency and precision of hallucination detection. Furthermore, it requires fewer retrieval steps on average, thus decreasing response times 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- HaloScope: Harnessing Unlabeled LLM Generations for Hallucination DetectionXuefeng Du, Chaowei Xiao, Sharon LiNeurIPS 2024 · 131 citations
- Searching for Best Practices in Retrieval-Augmented GenerationXiaohua Wang, Zhenghua Wang, Xuan Gao, Feiran Zhang et al.EMNLP 2024 · 75 citations
- Enhancing Uncertainty Modeling with Semantic Graph for Hallucination DetectionKedi Chen, Qin Chen, Jie Zhou, Xinqi Tao et al.AAAI 2025 · 13 citations
- Citation-Enhanced Generation for LLM-based ChatbotsWeitao Li, Junkai Li, Weizhi Ma, Yang LiuACL 2024 · 12 citations
- ELDER: Enhancing Lifelong Model Editing with Mixture-of-LoRAJiaang Li, Quan Wang, Zhongnan Wang, Yongdong Zhang et al.AAAI 2025 · 6 citations
Builds on7
- Adversarial NLI: A New Benchmark for Natural Language UnderstandingYixin Nie, Adina Williams, Emily Dinan, Mohit Bansal et al.ACL 2020 · 602 citations
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding SharingPengcheng He, Jianfeng Gao, Weizhu ChenICLR 2023 · 394 citations
- SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language ModelsPotsawee Manakul, Adian Liusie, Mark J. F. GalesEMNLP 2023 · 331 citations
- Factuality Enhanced Language Models for Open-Ended Text GenerationNayeon Lee, Wei Ping, Peng Xu, Mostofa Patwary et al.NeurIPS 2022 · 318 citations
- Self-contradictory Hallucinations of Large Language Models: Evaluation, Detection and MitigationNiels Mündler, Jingxuan He, Slobodan Jenko, Martin T. VechevICLR 2024 · 172 citations
Related papers
- Efficient Hallucination Detection: Adaptive Bayesian Estimation of Semantic Entropy with Guided Semantic ExplorationQiyao Sun, Xingming Li, Xixiang He, Ao Cheng et al.AAAI 2026 · 1 citation
- InterrogateLLM: Zero-Resource Hallucination Detection in LLM-Generated AnswersYakir Yehuda, Itzik Malkiel, Oren Barkan, Jonathan Weill et al.ACL 2024
- Enhancing Hallucination Detection through Noise InjectionLitian Liu, Reza Pourreza, Sunny Panchal, Apratim Bhattacharyya et al.ICLR 2026 · 19 citations
- LAFaCT: Attribution-based Localization and Focused Sequential Analysis of Fact-Critical Tokens for Hallucination DetectionXin Wang, Jiahao Li, Licheng Zhang, Zhendong MaoACL 2026
- Why LVLMs are More Prone to Hallucinations in Longer Responses: The Role of ContextGe Zheng, Jiaye Qian, Jiajin Tang, Sibei YangICCV 2025 · 2 citations
