Enhancing Language Model Factuality via Activation-Based Confidence Calibration and Guided Decoding
Xin Liu, Farima Fatahi Bayat, Lu Wang
Abstract
Calibrating language models (LMs) aligns their generation confidence with the actual likelihood of answer correctness, which can inform users about LMs' reliability and mitigate hallucinated content. However, prior calibration methods, such as self-consistency-based and logit-based approaches, are either limited in inference-time efficiency or fall short of providing informative signals. Moreover, simply filtering out low-confidence responses reduces the LM's helpfulness when the answers are correct. Therefore, effectively using calibration techniques to enhance an LM's factuality remains an unsolved challenge. In this paper, we first propose an activation-based calibration method, ACTCAB, which trains a linear layer on top of the LM's last-layer activations that can better capture the representations of knowledge. Built on top of ACTCAB, we further propose CODEC, a confidence-guided decoding strategy to elicit truthful answers with high confidence from LMs. By evaluating on five popular QA benchmarks, ACTCAB achieves superior calibration performance than all competitive baselines, e.g., by reducing the average expected calibration error (ECE) score by up to 39%. Further experiments on CODEC show consistent improvements in several LMs' factuality on challenging QA datasets, such as TruthfulQA, highlighting the value of confidence signals in enhancing the factuality. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9b8070b8-7acb-4ec2-9dc5-a97ebbe0af27Cited by top-tier papers6
- Calibration and Correctness of Language Models for CodeClaudio Spiess, David Gros, Kunal Suresh Pai, Michael Pradel et al.ICSE 2025 · 21 citations
- Attention-guided Self-reflection for Zero-shot Hallucination Detection in Large Language ModelsQiang Liu, Xinlong Chen, Yue Ding, Bowen Song et al.EMNLP 2025 · 2 citations
- Confidence Should Be Calibrated More Than One Turn DeepZhaohan Zhang, Chengzhengxu Li, Xiaoming Liu, Chao Shen et al.ACL 2026 · 2 citations
- How Far Ahead Do LLMs Plan? Uncovering the Latent Horizon in Chain-of-Thought ReasoningLiyan Xu, Mo Yu, Fandong Meng, Jie ZhouICML 2026 · 1 citation
- GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language ModelsZhaohan Zhang, Ziquan Liu, Ioannis PatrasACL 2026
Builds on8
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister et al.NeurIPS 2023 · 1,549 citations
- Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMsMiao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li et al.ICLR 2024 · 867 citations
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
- Prompting GPT-3 To Be ReliableChenglei Si, Zhe Gan, Zhengyuan Yang, Shuohang Wang et al.ICLR 2023 · 68 citations
Related papers
- In-Context Sharpness as Alerts: An Inner Representation Perspective for Hallucination MitigationShiqi Chen, Miao Xiong, Junteng Liu, Zhengxuan Wu et al.ICML 2024 · 49 citations
- LoVeC: Reinforcement Learning for Better Verbalized Confidence in Long-Form GenerationCaiqi Zhang, Xiaochen Zhu, Chengzu Li, Nigel Collier et al.ACL 2026 · 16 citations
- LitCab: Lightweight Language Model Calibration over Short- and Long-form ResponsesXin Liu, Muhammad Khalifa, Lu WangICLR 2024 · 41 citations
- Characterizing Truthfulness in Large Language Model Generations with Local Intrinsic DimensionFan Yin, Jayanth Srinivasa, Kai-Wei ChangICML 2024 · 43 citations
- Adaptive Activation Steering: A Tuning-Free LLM Truthfulness Improvement Method for Diverse Hallucinations CategoriesTianlong Wang, Xianfeng Jiao, Yinghao Zhu, Zhongzhi Chen et al.WWW 2025 · 64 citations
