HaloScope: Harnessing Unlabeled LLM Generations for Hallucination Detection
Xuefeng Du, Chaowei Xiao, Sharon Li
Abstract
The surge in applications of large language models (LLMs) has prompted concerns about the generation of misleading or fabricated information, known as hallucinations. Therefore, detecting hallucinations has become critical to maintaining trust in LLM-generated content. A primary challenge in learning a truthfulness classifier is the lack of a large amount of labeled truthful and hallucinated data. To address the challenge, we introduce HaloScope, a novel learning framework that leverages the unlabeled LLM generations in the wild for hallucination detection. Such unlabeled data arises freely upon deploying LLMs in the open world, and consists of both truthful and hallucinated information. To harness the unlabeled data, we present an automated membership estimation score for distinguishing between truthful and untruthful generations within unlabeled mixture data, thereby enabling the training of a binary truthfulness classifier on top. Importantly, our framework does not require extra data collection and human annotations, offering strong flexibility and practicality for real-world applications. Extensive experiments show that HaloScope can achieve superior hallucination detection performance, outperforming the competitive rivals by a significant margin. Code is available at https://github.com/deeplearningwisc/haloscope.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4a303a78-d20b-423b-92fb-e8241e85c6e4Cited by top-tier papers25
- SAFE: Multitask Failure Detection for Vision-Language-Action ModelsQiao Gu, Yuanliang Ju, Shengxiang Sun, Igor Gilitschenski et al.NeurIPS 2025 · 103 citations
- Robust Hallucination Detection in LLMs via Adaptive Token SelectionMengjia Niu, Hamed Haddadi, Guansong PangNeurIPS 2025 · 24 citations
- Auditing Meta-Cognitive Hallucinations in Reasoning Large Language ModelsHaolang Lu, Yilian Liu, Jingxin Xu, Guoshun Nan et al.NeurIPS 2025 · 23 citations
- Hallucination Detection in LLMs with Topological Divergence on Attention GraphsAlexandra Bazarova, Andrei Volodichev, Aleksandr Yugay, Andrey Shulga et al.ACL 2026 · 14 citations
- GLSim: Detecting Object Hallucinations in LVLMs via Global-Local SimilaritySeongheon Park, Sharon LiNeurIPS 2025 · 13 citations
Builds on20
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister et al.NeurIPS 2023 · 1,549 citations
- Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMsMiao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li et al.ICLR 2024 · 867 citations
- Uncertainty Estimation in Autoregressive Structured PredictionAndrey Malinin, Mark J. F. GalesICLR 2021 · 439 citations
- DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language ModelsYung-Sung Chuang, Yujia Xie, Hongyin Luo, Yoon Kim et al.ICLR 2024 · 354 citations
Related papers
- The Dawn After the Dark: An Empirical Study on Factuality Hallucination in Large Language ModelsJunyi Li, Jie Chen, Ruiyang Ren, Xiaoxue Cheng et al.ACL 2024 · 49 citations
- Unsupervised Hallucination Detection by Inspecting Reasoning ProcessesPonhvoan Srey, Xiaobao Wu, Anh Tuan LuuEMNLP 2025
- The Digital Dunning-Kruger Effect: Decoupling Hallucinations via Geometric Hidden-state Observation for Semantic TruthfulnessYueheng Mao, Min Yu, Gengwang Li, Jianguo Jiang et al.ACL 2026
- HD-NDEs: Neural Differential Equations for Hallucination Detection in LLMsQing Li, Jiahui Geng, Zongxiong Chen, Derui Zhu et al.ACL 2025
- HALoGEN: Fantastic LLM Hallucinations and Where to Find ThemAbhilasha Ravichander, Shrusti Ghela, David Wadden, Yejin ChoiACL 2025 · 35 citations
