HalluLens: LLM Hallucination Benchmark
Yejin Bang, Ziwei Ji, Alan Schelten, Anthony Hartshorn, Tara Fowler, Cheng Zhang, Nicola Cancedda, Pascale Fung
摘要
Large language models (LLMs) often generate responses that deviate from user input or training data, a phenomenon known as "hallucination." These hallucinations undermine user trust and hinder the adoption of generative AI systems. Addressing hallucinations is important for the advancement of LLMs. This paper introduces a comprehensive hallucination benchmark HalluLens, incorporating both extrinsic and intrinsic evaluation tasks, built upon a clear taxonomy of hallucination. A major challenge in benchmarking hallucinations is the lack of a unified framework due to inconsistent definitions and categorizations. We disentangle LLM hallucination from "factuality" and propose a taxonomy distinguishing extrinsic and intrinsic hallucinations to promote consistency and facilitate research. We emphasize extrinsic hallucinations -where generated content deviates from training data -as they become increasingly relevant with LLM advancements. However, no benchmark is solely dedicated to extrinsic hallucinations. To address this gap, HalluLens introduces three new extrinsic tasks with dynamic test set generation to mitigate data leakage and ensure robustness. We release codebase for extrinsic hallucination benchmark. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Beyond Prompt-Induced Lies: Investigating LLM Deception on Benign PromptsZhaomin Wu, Mingzhe Du, See-Kiong Ng, Bingsheng HeICLR 2026 · 被引用 11 次
- Pixels Don't Lie (But Your Detector Might): Bootstrapping MLLM-as-a-Judge for Trustworthy Deepfake Detection and Reasoning SupervisionKartik Kuckreja, Parul Gupta, Muhammad Haris Khan, Abhinav DhallCVPR 2026 · 被引用 5 次
- HalluAudio: A Comprehensive Benchmark for Hallucination Detection in Large Audio-Language ModelsFeiyu Zhao, Yiming Chen, Wenhuan Lu, Daipeng Zhang 等ACL 2026 · 被引用 3 次
- HalluGen: Synthesizing Realistic and Controllable Hallucinations for Evaluating Image RestorationSeunghoi Kim, Henry F. J. Tregidgo, Chen Jin, Matteo Figini 等CVPR 2026 · 被引用 2 次
- Human-AI Interaction for Time-Critical Sensemaking in Missing Persons InvestigationsPola Zuzanna Labedzka, Dorian Peters, John J. Dudley, Miri ZilkaCHI 2026 · 被引用 1 次
它引用的顶会 Paper14
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 被引用 3,228 次
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski 等USENIX Security 2021 · 被引用 2,866 次
- SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language ModelsPotsawee Manakul, Adian Liusie, Mark J. F. GalesEMNLP 2023 · 被引用 331 次
- When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric MemoriesAlex Mallen, Akari Asai, Victor Zhong, Rajarshi Das 等ACL 2023 · 被引用 233 次
- FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text GenerationSewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis 等EMNLP 2023 · 被引用 225 次
相关 Paper
- CodeHalu: Investigating Code Hallucinations in LLMs via Execution-based VerificationYuchen Tian, Weixiang Yan, Qian Yang, Xuandong Zhao 等AAAI 2025 · 被引用 41 次
- HALoGEN: Fantastic LLM Hallucinations and Where to Find ThemAbhilasha Ravichander, Shrusti Ghela, David Wadden, Yejin ChoiACL 2025 · 被引用 35 次
- Hal-Eval: A Universal and Fine-grained Hallucination Evaluation Framework for Large Vision Language ModelsChaoya Jiang, Hongrui Jia, Mengfan Dong, Wei Ye 等ACM MM 2024 · 被引用 19 次
- UHGEval: Benchmarking the Hallucination of Chinese Large Language Models via Unconstrained GenerationXun Liang, Shichao Song, Simin Niu, Zhiyu Li 等ACL 2024 · 被引用 15 次
- The Dawn After the Dark: An Empirical Study on Factuality Hallucination in Large Language ModelsJunyi Li, Jie Chen, Ruiyang Ren, Xiaoxue Cheng 等ACL 2024 · 被引用 49 次
