From Flat Facts to Sharp Hallucinations: Detecting Stubborn Errors via Gradient Sensitivity
Liew Yee Zhing, Andrew Tan, Anwar Majeed
摘要
Traditional hallucination detection fails on "Stubborn Hallucinations" — errors where LLMs are confidently wrong. We propose a geometric solution: Embedding-Perturbed Gradient Sensitivity (EPGS). We hypothesize that while robust facts reside in flat minima, stubborn hallucinations sit in sharp minima, supported by brittle memorization. EPGS detects this sharpness by perturbing input embeddings with Gaussian noise and measuring the resulting spike in gradient magnitude. This acts as an efficient proxy for the Hessian spectrum, differentiating stable knowledge from unstable memorization. Our experiments show that EPGS significantly outperforms entropy-based and representation-based baselines, providing a robust signal for detecting high-confidence factual errors.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper3
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan 等ICLR 2020 · 被引用 705 次
- Uncertainty Estimation in Autoregressive Structured PredictionAndrey Malinin, Mark J. F. GalesICLR 2021 · 被引用 439 次
- HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language ModelsJunyi Li, Xiaoxue Cheng, Xin Zhao, Jian-Yun Nie 等EMNLP 2023 · 被引用 224 次
相关 Paper
- Lyapunov Probes for Hallucination Detection in Large Foundation ModelsBozhi Luan, Gen Li, Yalan Qin, Jifeng Guo 等CVPR 2026 · 被引用 3 次
- The Digital Dunning-Kruger Effect: Decoupling Hallucinations via Geometric Hidden-state Observation for Semantic TruthfulnessYueheng Mao, Min Yu, Gengwang Li, Jianguo Jiang 等ACL 2026
- A Geometric Analysis of Small-sized Language Model HallucinationsEmanuele Ricco, Elia Onofri, Lorenzo Cima, Stefano Cresci 等ICML 2026 · 被引用 1 次
- REMIND: Memorization and Unlearning in LLMs Through the Lens of Input Loss LandscapesLiran Cohen, Yaniv Nemcovsky, Avi MendelsonACL 2026
- Hallucination Detox: Sensitivity Dropout (SenD) for Large Language Model TrainingShahrad Mohammadzadeh, Juan David Guerra, Marco Bonizzato, Reihaneh Rabbany 等ACL 2025 · 被引用 4 次
