HALoGEN: Fantastic LLM Hallucinations and Where to Find Them
Abhilasha Ravichander, Shrusti Ghela, David Wadden, Yejin Choi
摘要
Despite their impressive ability to generate high-quality and fluent text, generative large language models (LLMs) also produce hallucinations: statements that are misaligned with established world knowledge or provided input context. However, measuring hallucination can be challenging, as having humans verify model generations on-the-fly is both expensive and time-consuming. In this work, we release HALOGEN , a comprehensive hallucination benchmark consisting of: (1) 10,923 prompts for generative models spanning nine domains including programming, scientific attribution, and summarization, and (2) automatic highprecision verifiers for each use case that decompose LLM generations into atomic units, and verify each unit against a high-quality knowledge source. We use this framework to evaluate ∼150,000 generations from 14 language models, finding that even the best-performing models are riddled with hallucinations (sometimes up to 86% of generated atomic facts depending on the domain). We further define a novel error classification for LLM hallucinations based on whether they likely stem from incorrect recollection of training data (Type A errors), or incorrect knowledge in training data (Type B errors), or are fabrication (Type C errors). We hope our framework provides a foundation to enable the principled study of why generative models hallucinate, and advances the development of trustworthy large language models. * Equal Contribution † Independent researcher, work done in part while author was at the University of Washington.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Why and How LLMs Hallucinate: Connecting the Dots with Subsequence AssociationsYiyou Sun, Yu Gai, Lijie Chen, Abhilasha Ravichander 等NeurIPS 2025 · 被引用 20 次
- Empty Shelves or Lost Keys? Recall Is the Bottleneck for Parametric FactualityNitay Calderon, Eyal Ben-David, Zorik Gekhman, Eran Ofek 等ICML 2026 · 被引用 8 次
- Teaching Large Language Models to Maintain Contextual Faithfulness via Synthetic Tasks and Reinforcement LearningShuzheng Si, Haozhe Zhao, Cheng Gao, Yuzhuo Bai 等AAAI 2026 · 被引用 4 次
- OWL: Probing Cross-Lingual Recall of Memorized Texts via World LiteratureAlisha Srivastava, Emir Korukluoglu, Minh Nhat Le, Duyen Tran 等EMNLP 2025 · 被引用 1 次
- The Algorithmic Self-Portrait: Deconstructing Memory in ChatGPTAbhisek Dash, Soumi Das, Elisabeth Kirsten, Qinyuan Wu 等WWW 2026 · 被引用 1 次
它引用的顶会 Paper14
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 被引用 3,228 次
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
- How Language Model Hallucinations Can SnowballMuru Zhang, Ofir Press, William Merrill, Alisa Liu 等ICML 2024 · 被引用 406 次
- SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language ModelsPotsawee Manakul, Adian Liusie, Mark J. F. GalesEMNLP 2023 · 被引用 331 次
相关 Paper
- CodeHalu: Investigating Code Hallucinations in LLMs via Execution-based VerificationYuchen Tian, Weixiang Yan, Qian Yang, Xuandong Zhao 等AAAI 2025 · 被引用 41 次
- HalluLens: LLM Hallucination BenchmarkYejin Bang, Ziwei Ji, Alan Schelten, Anthony Hartshorn 等ACL 2025
- The Dawn After the Dark: An Empirical Study on Factuality Hallucination in Large Language ModelsJunyi Li, Jie Chen, Ruiyang Ren, Xiaoxue Cheng 等ACL 2024 · 被引用 49 次
- K-HALU: Multiple Answer Korean Hallucination Benchmark for Large Language ModelsJaehyung Seo, Heuiseok LimICLR 2025
- UHGEval: Benchmarking the Hallucination of Chinese Large Language Models via Unconstrained GenerationXun Liang, Shichao Song, Simin Niu, Zhiyu Li 等ACL 2024 · 被引用 15 次
