The Illusion of Progress: Re-evaluating Hallucination Detection in LLMs
Denis Janiak, Jakub Binkowski, Albert Sawczyn, Bogdan Gabrys, Ravid Shwartz-Ziv, Tomasz Kajdanowicz
摘要
Large language models (LLMs) have revolutionized natural language processing, yet their tendency to hallucinate poses serious challenges for reliable deployment. Despite numerous hallucination detection methods, their evaluations often rely on ROUGE, a metric based on lexical overlap that misaligns with human judgments. Through comprehensive human studies, we demonstrate that while ROUGE exhibits high recall, its extremely low precision leads to misleading performance estimates. In fact, several established detection methods show performance drops of up to 45.9% when assessed using human-aligned metrics like LLM-as-Judge. Moreover, our analysis reveals that simple heuristics based on response length can rival complex detection techniques, exposing a fundamental flaw in current evaluation practices. We argue that adopting semantically aware and robust evaluation frameworks is essential to accurately gauge the true performance of hallucination detection methods, ultimately ensuring the trustworthiness of LLM outputs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- HalluGuard: Demystifying Data-Driven and Reasoning-Driven Hallucinations in LLMsXinyue Zeng, Junhong Lin, Yujun Yan, Feng Guo 等ICLR 2026 · 被引用 13 次
- Beyond In-Domain Detection: SpikeScore for Cross-Domain Hallucination DetectionYongxin Deng, Zhen Fang, Sharon Li, Ling ChenICLR 2026 · 被引用 5 次
- Knowing When to Quit: A Principled Framework for Dynamic Abstention in LLM ReasoningHen Davidov, Nachshon Cohen, Oren Kalinsky, Yaron Fairstein 等ICML 2026 · 被引用 3 次
- Automatic Layer Selection for Hallucination DetectionXinpeng Wang, William Cao, Andrew Wilson, Zhe ZengICML 2026 · 被引用 2 次
- Dust Off Kindle Highlights With Quologue: Surfacing Personal Data With Generative AI for Reflective ExperiencesSol Kang, William Odom, Amy Yo Sue Chen, Carman NeustaedterCHI 2026 · 被引用 1 次
它引用的顶会 Paper14
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Uncertainty Estimation in Autoregressive Structured PredictionAndrey Malinin, Mark J. F. GalesICLR 2021 · 被引用 439 次
- How Language Model Hallucinations Can SnowballMuru Zhang, Ofir Press, William Merrill, Alisa Liu 等ICML 2024 · 被引用 406 次
- INSIDE: LLMs' Internal States Retain the Power of Hallucination DetectionChao Chen, Kai Liu, Ze Chen, Yi Gu 等ICLR 2024 · 被引用 281 次
- HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language ModelsJunyi Li, Xiaoxue Cheng, Xin Zhao, Jian-Yun Nie 等EMNLP 2023 · 被引用 224 次
相关 Paper
- FineSurE: Fine-grained Summarization Evaluation using LLMsHwanjun Song, Hang Su, Igor Shalyminov, Jason Cai 等ACL 2024
- The Dawn After the Dark: An Empirical Study on Factuality Hallucination in Large Language ModelsJunyi Li, Jie Chen, Ruiyang Ren, Xiaoxue Cheng 等ACL 2024 · 被引用 49 次
- Estimating LLM Consistency: A User Baseline vs Surrogate MetricsXiaoyuan Wu, Weiran Lin, Omer Akgul, Lujo BauerEMNLP 2025
- Identifying Reliable Evaluation Metrics for Scientific Text RevisionLéane Jourdan, Nicolas Hernandez, Florian Boudin, Richard DufourACL 2025
- Addressing Pitfalls in the Evaluation of Uncertainty Estimation Methods for Natural Language GenerationMykyta Ielanskyi, Kajetan Schweighofer, Lukas Aichberger, Sepp HochreiterICLR 2026 · 被引用 10 次
