HALoGEN: Fantastic LLM Hallucinations and Where to Find Them
Abhilasha Ravichander, Shrusti Ghela, David Wadden, Yejin Choi
Abstract
Despite their impressive ability to generate high-quality and fluent text, generative large language models (LLMs) also produce hallucinations: statements that are misaligned with established world knowledge or provided input context. However, measuring hallucination can be challenging, as having humans verify model generations on-the-fly is both expensive and time-consuming. In this work, we release HALOGEN , a comprehensive hallucination benchmark consisting of: (1) 10,923 prompts for generative models spanning nine domains including programming, scientific attribution, and summarization, and (2) automatic highprecision verifiers for each use case that decompose LLM generations into atomic units, and verify each unit against a high-quality knowledge source. We use this framework to evaluate ∼150,000 generations from 14 language models, finding that even the best-performing models are riddled with hallucinations (sometimes up to 86% of generated atomic facts depending on the domain). We further define a novel error classification for LLM hallucinations based on whether they likely stem from incorrect recollection of training data (Type A errors), or incorrect knowledge in training data (Type B errors), or are fabrication (Type C errors). We hope our framework provides a foundation to enable the principled study of why generative models hallucinate, and advances the development of trustworthy large language models. * Equal Contribution † Independent researcher, work done in part while author was at the University of Washington.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f9e6eacf-1ac5-4dea-9a5f-f7266f2aca15Cited by top-tier papers8
- Why and How LLMs Hallucinate: Connecting the Dots with Subsequence AssociationsYiyou Sun, Yu Gai, Lijie Chen, Abhilasha Ravichander et al.NeurIPS 2025 · 20 citations
- Empty Shelves or Lost Keys? Recall Is the Bottleneck for Parametric FactualityNitay Calderon, Eyal Ben-David, Zorik Gekhman, Eran Ofek et al.ICML 2026 · 8 citations
- Teaching Large Language Models to Maintain Contextual Faithfulness via Synthetic Tasks and Reinforcement LearningShuzheng Si, Haozhe Zhao, Cheng Gao, Yuzhuo Bai et al.AAAI 2026 · 4 citations
- OWL: Probing Cross-Lingual Recall of Memorized Texts via World LiteratureAlisha Srivastava, Emir Korukluoglu, Minh Nhat Le, Duyen Tran et al.EMNLP 2025 · 1 citation
- The Algorithmic Self-Portrait: Deconstructing Memory in ChatGPTAbhisek Dash, Soumi Das, Elisabeth Kirsten, Qinyuan Wu et al.WWW 2026 · 1 citation
Builds on14
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- How Language Model Hallucinations Can SnowballMuru Zhang, Ofir Press, William Merrill, Alisa Liu et al.ICML 2024 · 406 citations
- SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language ModelsPotsawee Manakul, Adian Liusie, Mark J. F. GalesEMNLP 2023 · 331 citations
Related papers
- CodeHalu: Investigating Code Hallucinations in LLMs via Execution-based VerificationYuchen Tian, Weixiang Yan, Qian Yang, Xuandong Zhao et al.AAAI 2025 · 41 citations
- HalluLens: LLM Hallucination BenchmarkYejin Bang, Ziwei Ji, Alan Schelten, Anthony Hartshorn et al.ACL 2025
- The Dawn After the Dark: An Empirical Study on Factuality Hallucination in Large Language ModelsJunyi Li, Jie Chen, Ruiyang Ren, Xiaoxue Cheng et al.ACL 2024 · 49 citations
- K-HALU: Multiple Answer Korean Hallucination Benchmark for Large Language ModelsJaehyung Seo, Heuiseok LimICLR 2025
- UHGEval: Benchmarking the Hallucination of Chinese Large Language Models via Unconstrained GenerationXun Liang, Shichao Song, Simin Niu, Zhiyu Li et al.ACL 2024 · 15 citations
