LitCab: Lightweight Language Model Calibration over Short- and Long-form Responses
Xin Liu, Muhammad Khalifa, Lu Wang
Abstract
A model is considered well-calibrated when its probability estimate aligns with the actual likelihood of the output being correct. Calibrating language models (LMs) is crucial, as it plays a vital role in detecting and mitigating hallucinations of LMs as well as building more trustworthy models. However, standard calibration techniques may not be suited for LM calibration. For instance, post-processing methods such as temperature scaling do not reorder the candidate generations. On the other hand, training-based methods require fine-tuning the entire model, which is impractical for LMs of large scale. We present LITCAB, a lightweight calibration mechanism consisting of a single linear layer that takes the input text representation and predicts a bias term, which is then added to the LM output logits. LITCAB improves model calibration by only adding < 2% of the original model parameters. For evaluation, we construct CAT, a benchmark consisting of eight text generation tasks, covering responses ranging from short phrases to paragraphs. We test LITCAB with Llama2-7B, where it improves calibration across all tasks, reducing the average ECE score by as large as 30%. We further conduct a comprehensive evaluation with multiple popular open-sourced LMs from GPT and LLaMA families, yielding the following key findings: (i) Larger models within the same family exhibit better calibration on tasks with short generation tasks, but not necessarily for longer ones. (ii) GPT-family models show superior calibration compared to LLaMA, Llama2, and Vicuna models, despite having much fewer parameters. (iii) Fine-tuning pretrained model (e.g., LLaMA) with samples of limited purpose (e.g., conversations) may lead to worse calibration, highlighting the importance of fine-tuning setups for calibrating LMs. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 303e7b16-af43-4b05-b555-c37e651530cbCited by top-tier papers11
- Your Pre-trained LLM is Secretly an Unsupervised Confidence CalibratorBeier Luo, Shuoyuan Wang, Sharon Li, Hongxin WeiNeurIPS 2025 · 22 citations
- Calibration and Correctness of Language Models for CodeClaudio Spiess, David Gros, Kunal Suresh Pai, Michael Pradel et al.ICSE 2025 · 21 citations
- LoVeC: Reinforcement Learning for Better Verbalized Confidence in Long-Form GenerationCaiqi Zhang, Xiaochen Zhu, Chengzu Li, Nigel Collier et al.ACL 2026 · 16 citations
- Generalized Correctness Models: Learning Calibrated and Cross-Model Correctness Predictors from Historical PatternsHanqi Xiao, Vaidehi Patil, Hyunji Lee, Elias Stengel-Eskin et al.ICML 2026 · 5 citations
- Boosting Resilience of Large Language Models through Causality-Driven Robust OptimizationXiaoling Zhou, Mingjie Zhang, Zhemg Lee, Yuncheng Hua et al.NeurIPS 2025 · 5 citations
Builds on7
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
- FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text GenerationSewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis et al.EMNLP 2023 · 225 citations
- Large Dual Encoders Are Generalizable RetrieversJianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai et al.EMNLP 2022 · 145 citations
Related papers
- Enhancing Language Model Factuality via Activation-Based Confidence Calibration and Guided DecodingXin Liu, Farima Fatahi Bayat, Lu WangEMNLP 2024 · 1 citation
- Thermometer: Towards Universal Calibration for Large Language ModelsMaohao Shen, Subhro Das, Kristjan H. Greenewald, Prasanna Sattigeri et al.ICML 2024 · 38 citations
- Towards Objective Fine-tuning: How LLMs' Prior Knowledge Causes Potential Poor Calibration?Ziming Wang, Zeyu Shi, Haoyi Zhou, Shiqi Gao et al.ACL 2025 · 6 citations
- Linguistic Calibration of Long-Form GenerationsNeil Band, Xuechen Li, Tengyu Ma, Tatsunori HashimotoICML 2024 · 56 citations
- LACIE: Listener-Aware Finetuning for Calibration in Large Language ModelsElias Stengel-Eskin, Peter Hase, Mohit BansalNeurIPS 2024 · 26 citations
