Calibrated Language Model Fine-Tuning for In- and Out-of-Distribution Data
Lingkai Kong, Haoming Jiang, Yuchen Zhuang, Jie Lyu, Tuo Zhao, Chao Zhang
Abstract
Fine-tuned pre-trained language models can suffer from severe miscalibration for both in-distribution and out-of-distribution (OOD) data due to over-parameterization. To mitigate this issue, we propose a regularized fine-tuning method. Our method introduces two types of regularization for better calibration: (1) On-manifold regularization, which generates pseudo on-manifold samples through interpolation within the data manifold. Augmented training with these pseudo samples imposes a smoothness regularization to improve in-distribution calibration. (2) Off-manifold regularization, which encourages the model to output uniform distributions for pseudo off-manifold samples to address the over-confidence issue for OOD data. Our experiments demonstrate that the proposed method outperforms existing calibration methods for text classification in terms of expectation calibration error, misclassification detection, and OOD detection on six datasets. Our code can be found at https://github.com/Lingkai-Kong/ Calibrated-BERT-Fine-Tuning .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 87dc384c-e63b-42e4-ad0f-0fc1257fa336Cited by top-tier papers22
- On the Calibration of Pre-trained Language Models using Mixup Guided by Area Under the Margin and SaliencySeoyeon Park, Cornelia CarageaACL 2022 · 44 citations
- Active Prompting with Chain-of-Thought for Large Language ModelsShizhe Diao, Pengcheng Wang, Yong Lin, Rui Pan et al.ACL 2024 · 40 citations
- Don't Hallucinate, Abstain: Identifying LLM Knowledge Gaps via Multi-LLM CollaborationShangbin Feng, Weijia Shi, Yike Wang, Wenxuan Ding et al.ACL 2024 · 30 citations
- Navigating the Grey Area: How Expressions of Uncertainty and Overconfidence Affect Language ModelsKaitlyn Zhou, Dan Jurafsky, Tatsunori HashimotoEMNLP 2023 · 29 citations
- When in Doubt: Neural Non-Parametric Uncertainty Quantification for Epidemic ForecastingHarshavardhan Kamarthi, Lingkai Kong, Alexander Rodríguez, Chao Zhang et al.NeurIPS 2021 · 26 citations
Builds on3
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 1,352 citations
- SMART: Robust and Efficient Fine-Tuning for Pre-trained Natural Language Models through Principled Regularized OptimizationHaoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu et al.ACL 2020 · 148 citations
Related papers
- OoMMix: Out-of-manifold Regularization in Contextual Embedding Space for Text ClassificationSeonghyeon Lee, Dongha Lee, Hwanjo YuACL 2021
- MASKER: Masked Keyword Regularization for Reliable Text ClassificationSeung Jun Moon, Sangwoo Mo, Kimin Lee, Jaeho Lee et al.AAAI 2021 · 39 citations
- Understanding and Mitigating Miscalibration in Prompt Tuning for Vision-Language ModelsShuoyuan Wang, Yixuan Li, Hongxin WeiICML 2025
- Preserving Pre-trained Features Helps Calibrate Fine-tuned Language ModelsGuande He, Jianfei Chen, Jun ZhuICLR 2023 · 1 citation
- Restoring Calibration for Aligned Large Language Models: A Calibration-Aware Fine-Tuning ApproachJiancong Xiao, Bojian Hou, Zhanliang Wang, Ruochen Jin et al.ICML 2025
