HaDeMiF: Hallucination Detection and Mitigation in Large Language Models
Xiaoling Zhou, Mingjie Zhang, Zhemg Lee, Wei Ye, Shikun Zhang
摘要
The phenomenon of knowledge hallucinations has raised substantial concerns about the security and reliability of deployed large language models (LLMs). Current methods for detecting hallucinations primarily depend on manually designed individual metrics, such as prediction uncertainty and consistency, and fall short in effectively calibrating model predictions, thus constraining their detection accuracy and applicability in practical applications. In response, we propose an advanced framework, termed HADEMIF, for detecting and mitigating hallucinations in LLMs. Specifically, hallucinations within the output and semantic spaces of LLMs are comprehensively captured through two compact networks-a novel, interpretable tree model known as the Deep Dynamic Decision Tree (D3T) and a Multilayer Perceptron (MLP)-which take as input a set of prediction characteristics and the hidden states of tokens, respectively. The predictions of LLMs are subsequently calibrated using the outputs from the D3T and MLP networks, aiming to mitigate hallucinations and enhance model calibration. HADEMIF can be applied during both the inference and fine-tuning phases of LLMs, introducing less than 2% of the parameters relative to the LLMs through the training of two small-scale networks. Extensive experiments conclusively demonstrate the effectiveness of our framework in hallucination detection and model calibration across text generation tasks with responses of varying lengths.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- PriorDrive: Enhancing Online HD Mapping with Unified Vector PriorsShuang Zeng, Xinyuan Chang, Xinran Liu, Yujian Yuan 等AAAI 2026 · 被引用 12 次
- Toward Faithful Retrieval-Augmented Generation with Sparse AutoencodersGuangzhi Xiong, Zhenghao He, Bohan Liu, Sanchit Sinha 等ICLR 2026 · 被引用 8 次
- Generalized Correctness Models: Learning Calibrated and Cross-Model Correctness Predictors from Historical PatternsHanqi Xiao, Vaidehi Patil, Hyunji Lee, Elias Stengel-Eskin 等ICML 2026 · 被引用 5 次
- Boosting Resilience of Large Language Models through Causality-Driven Robust OptimizationXiaoling Zhou, Mingjie Zhang, Zhemg Lee, Yuncheng Hua 等NeurIPS 2025 · 被引用 5 次
- Neural-Driven Image EditingPengfei Zhou, Jie Xia, Xiaopeng Peng, Wangbo Zhao 等NeurIPS 2025 · 被引用 5 次
它引用的顶会 Paper28
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMsMiao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li 等ICLR 2024 · 被引用 867 次
- Are Emergent Abilities of Large Language Models a Mirage?Rylan Schaeffer, Brando Miranda, Sanmi KoyejoNeurIPS 2023 · 被引用 796 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
- Uncertainty Estimation in Autoregressive Structured PredictionAndrey Malinin, Mark J. F. GalesICLR 2021 · 被引用 439 次
相关 Paper
- The Dawn After the Dark: An Empirical Study on Factuality Hallucination in Large Language ModelsJunyi Li, Jie Chen, Ruiyang Ren, Xiaoxue Cheng 等ACL 2024 · 被引用 49 次
- I Don't Know: Explicit Modeling of Uncertainty with an [IDK] TokenRoi Cohen, Konstantin Dobler, Eden Biran, Gerard de MeloNeurIPS 2024 · 被引用 35 次
- HD-NDEs: Neural Differential Equations for Hallucination Detection in LLMsQing Li, Jiahui Geng, Zongxiong Chen, Derui Zhu 等ACL 2025
- Enhancing Uncertainty Modeling with Semantic Graph for Hallucination DetectionKedi Chen, Qin Chen, Jie Zhou, Xinqi Tao 等AAAI 2025 · 被引用 13 次
- HARP: Hallucination Detection via Reasoning Subspace ProjectionJunjie Hu, Gang Tu, Shengyu Cheng, Jinxin Li 等ICLR 2026 · 被引用 6 次
