From Flat Facts to Sharp Hallucinations: Detecting Stubborn Errors via Gradient Sensitivity
Liew Yee Zhing, Andrew Tan, Anwar Majeed
Abstract
Traditional hallucination detection fails on "Stubborn Hallucinations" — errors where LLMs are confidently wrong. We propose a geometric solution: Embedding-Perturbed Gradient Sensitivity (EPGS). We hypothesize that while robust facts reside in flat minima, stubborn hallucinations sit in sharp minima, supported by brittle memorization. EPGS detects this sharpness by perturbing input embeddings with Gaussian noise and measuring the resulting spike in gradient magnitude. This acts as an efficient proxy for the Hessian spectrum, differentiating stable knowledge from unstable memorization. Our experiments show that EPGS significantly outperforms entropy-based and representation-based baselines, providing a robust signal for detecting high-confidence factual errors.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 28605439-4bb1-4fa9-8d58-5f1161dc2a12Builds on3
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan et al.ICLR 2020 · 705 citations
- Uncertainty Estimation in Autoregressive Structured PredictionAndrey Malinin, Mark J. F. GalesICLR 2021 · 439 citations
- HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language ModelsJunyi Li, Xiaoxue Cheng, Xin Zhao, Jian-Yun Nie et al.EMNLP 2023 · 224 citations
Related papers
- Lyapunov Probes for Hallucination Detection in Large Foundation ModelsBozhi Luan, Gen Li, Yalan Qin, Jifeng Guo et al.CVPR 2026 · 3 citations
- The Digital Dunning-Kruger Effect: Decoupling Hallucinations via Geometric Hidden-state Observation for Semantic TruthfulnessYueheng Mao, Min Yu, Gengwang Li, Jianguo Jiang et al.ACL 2026
- A Geometric Analysis of Small-sized Language Model HallucinationsEmanuele Ricco, Elia Onofri, Lorenzo Cima, Stefano Cresci et al.ICML 2026 · 1 citation
- REMIND: Memorization and Unlearning in LLMs Through the Lens of Input Loss LandscapesLiran Cohen, Yaniv Nemcovsky, Avi MendelsonACL 2026
- Hallucination Detox: Sensitivity Dropout (SenD) for Large Language Model TrainingShahrad Mohammadzadeh, Juan David Guerra, Marco Bonizzato, Reihaneh Rabbany et al.ACL 2025 · 4 citations
