Transferable Post-hoc Calibration on Pretrained Transformers in Noisy Text Classification
Jun Zhang, Wen Yao, Xiaoqian Chen, Ling Feng
摘要
Recent work has demonstrated that pretrained transformers are overconfident in text classification tasks, which can be calibrated by the famous post-hoc calibration method temperature scaling (TS). Character or word spelling mistakes are frequently encountered in real applications and greatly threaten transformer model safety. Research on calibration under noisy settings is rare, and we focus on this direction. Based on a toy experiment, we discover that TS performs poorly when the datasets are perturbed by slight noise, such as swapping the characters, which results in distribution shift. We further utilize two metrics, predictive uncertainty and maximum mean discrepancy (MMD), to measure the distribution shift between clean and noisy datasets, based on which we propose a simple yet effective transferable TS method for calibrating models dynamically. To evaluate the performance of the proposed methods under noisy settings, we construct a benchmark consisting of four noise types and five shift intensities based on the QNLI, AG-News, and Emotion tasks. Experimental results on the noisy benchmark show that (1) the metrics are effective in measuring distribution shift and (2) transferable TS can significantly decrease the expected calibration error (ECE) compared with the competitive baseline ensemble TS by approximately 46.09%.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Do Large Language Models Know What They Are Capable Of?Casey O. Barkan, Sidney Black, Oliver SourbutICLR 2026 · 被引用 11 次
- DisCal: Distribution-Aware Calibration for Mathematical Reasoning Under Character-Level Noisy InputsBo Zhang, Jiawei Zhang, Cong Gao, Bingxu Han 等ACL 2026
- MLP Can Be a Good Transformer LearnerSihao Lin, Pumeng Lyu, Dongrui Liu, Tao Tang 等CVPR 2024
- Global-Local Confidence Fusion for Hallucination Detection in Mathematical Reasoning TaskBo Zhang, Cong Gao, Linkang Yang, Bingxu Han 等AAAI 2026
它引用的顶会 Paper4
- Mix-n-Match : Ensemble and Compositional Methods for Uncertainty Calibration in Deep LearningJize Zhang, Bhavya Kailkhura, Thomas Yong-Jin HanICML 2020 · 被引用 276 次
- Improving model calibration with accuracy versus uncertainty optimizationRanganath Krishnan, Omesh TickooNeurIPS 2020 · 被引用 217 次
- Soft Calibration Objectives for Neural NetworksArchit Karandikar, Nicholas Cain, Dustin Tran, Balaji Lakshminarayanan 等NeurIPS 2021 · 被引用 127 次
- ACT: an Attentive Convolutional Transformer for Efficient Text ClassificationPengfei Li, Peixiang Zhong, Kezhi Mao, Dongzhe Wang 等AAAI 2021 · 被引用 47 次
相关 Paper
- Calibrating Zero-shot Cross-lingual (Un-)structured PredictionsZhengping Jiang, Anqi Liu, Benjamin Van DurmeEMNLP 2022 · 被引用 4 次
- Robust Calibration with Multi-domain Temperature ScalingYaodong Yu, Stephen Bates, Yi Ma, Michael I. JordanNeurIPS 2022 · 被引用 58 次
- Rethinking Data Distillation: Do Not Overlook CalibrationDongyao Zhu, Yanbo Fang, Bowen Lei, Yiqun Xie 等ICCV 2023 · 被引用 19 次
- On the Limitations of Temperature Scaling for Distributions with OverlapsMuthu Chidambaram, Rong GeICLR 2024 · 被引用 11 次
- From Noisy Prediction to True Label: Noisy Prediction Calibration via Generative ModelHeeSun Bae, Seungjae Shin, Byeonghu Na, JoonHo Jang 等ICML 2022 · 被引用 29 次
