A Head to Predict and a Head to Question: Pre-trained Uncertainty Quantification Heads for Hallucination Detection in LLM Outputs
Artem Shelmanov, Ekaterina Fadeeva, Akim Tsvigun, Ivan Tsvigun, Zhuohan Xie, Igor Kiselev, Nico Daheim, Caiqi Zhang, Artem Vazhentsev, Mrinmaya Sachan, Preslav Nakov, Timothy Baldwin
Abstract
Large Language Models (LLMs) have the tendency to hallucinate, i.e., to sporadically generate false or fabricated information. This presents a major challenge, as hallucinations often appear highly convincing and users generally lack the tools to detect them. Uncertainty quantification (UQ) provides a framework for assessing the reliability of model outputs, aiding in the identification of potential hallucinations. In this work, we introduce pre-trained UQ heads: supervised auxiliary modules for LLMs that substantially enhance their ability to capture uncertainty compared to unsupervised UQ methods. Their strong performance stems from the transformer architecture in their design, in the form of informative features derived from LLM attention maps and logits. Our experiments show that these heads are highly robust and achieve state-of-the-art performance in claim-level hallucination detection across both in-domain and out-of-domain prompts. Moreover, these modules demonstrate strong generalization to languages they were not explicitly trained on. We pre-train a collection of UQ heads for popular LLM series, including Mistral, Llama, and Gemma. We publicly release both the code and the pre-trained heads. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6d62f90a-5f8f-4013-bb96-cf420d9e86a6Cited by top-tier papers3
- LoGU: Long-form Generation with Uncertainty ExpressionsRuihan Yang, Caiqi Zhang, Zhisong Zhang, Xinting Huang et al.ACL 2025 · 21 citations
- Efficient Test-Time Scaling of Multi-Step Reasoning by Probing Internal States of Large Language ModelsJingwei Ni, Ekaterina Fadeeva, Tianyi Wu, Mubashara Akhtar et al.ACL 2026 · 1 citation
- Learning Uncertainty from Sequential Internal Dispersion in Large Language ModelsPonhvoan Srey, Xiaobao Wu, Cong-Duy T. Nguyen, Anh Tuan LuuACL 2026
Builds on16
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- Uncertainty Estimation in Autoregressive Structured PredictionAndrey Malinin, Mark J. F. GalesICLR 2021 · 439 citations
- SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language ModelsPotsawee Manakul, Adian Liusie, Mark J. F. GalesEMNLP 2023 · 331 citations
- FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text GenerationSewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis et al.EMNLP 2023 · 225 citations
- Attention Satisfies: A Constraint-Satisfaction Lens on Factual Errors of Language ModelsMert Yüksekgönül, Varun Chandrasekaran, Erik Jones, Suriya Gunasekar et al.ICLR 2024 · 73 citations
Related papers
- Efficient Hallucination Detection for LLMs Using Uncertainty-Aware Attention HeadsArtem Vazhentsev, Lyudmila Rvanova, Gleb Kuzmin, Ekaterina Fadeeva et al.ICML 2026 · 16 citations
- Unconditional Truthfulness: Learning Unconditional Uncertainty of Large Language ModelsArtem Vazhentsev, Ekaterina Fadeeva, Rui Xing, Gleb Kuzmin et al.EMNLP 2025 · 1 citation
- Semantic Uncertainty Quantification of Hallucinations in LLMs: A Quantum Tensor Network Based MethodPragatheeswaran Vipulanandan, Kamal Premaratne, Dilip SarkarICLR 2026 · 4 citations
- Understanding and Mitigating Hallucination in Large Vision-Language Models via Modular Attribution and InterventionTianyun Yang, Ziniu Li, Juan Cao, Chang XuICLR 2025
- Hallucination Detection in Large Language Models with Metamorphic RelationsBorui Yang, Md Afif Al Mamun, Jie M. Zhang, Gias UddinFSE 2025 · 14 citations
