On the Calibration of Pre-trained Language Models using Mixup Guided by Area Under the Margin and Saliency
Seoyeon Park, Cornelia Caragea
Abstract
A well-calibrated neural model produces confidence (probability outputs) closely approximated by the expected accuracy. While prior studies have shown that mixup training as a data augmentation technique can improve model calibration on image classification tasks, little is known about using mixup for model calibration on natural language understanding (NLU) tasks. In this paper, we explore mixup for model calibration on several NLU tasks and propose a novel mixup strategy for pre-trained language models that improves model calibration further. Our proposed mixup is guided by both the Area Under the Margin (AUM) statistic (Pleiss et al., 2020) and the saliency map of each sample (Simonyan et al., 2013). Moreover, we combine our mixup strategy with model miscalibration correction techniques (i.e., label smoothing and temperature scaling) and provide detailed analyses of their impact on our proposed mixup. We focus on systematically designing experiments on three NLU tasks: natural language inference, paraphrase detection, and commonsense reasoning. Our method achieves the lowest expected calibration error compared to strong baselines on both in-domain and out-of-domain test samples while maintaining competitive accuracy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a56f8449-3468-4042-9365-73d43bd2e1daCited by top-tier papers12
- Bayesian Low-rank Adaptation for Large Language ModelsAdam X. Yang, Maxime Robeyns, Xi Wang, Laurence AitchisonICLR 2024 · 111 citations
- Thermometer: Towards Universal Calibration for Large Language ModelsMaohao Shen, Subhro Das, Kristjan H. Greenewald, Prasanna Sattigeri et al.ICML 2024 · 38 citations
- Spurious Feature Diversification Improves Out-of-distribution GeneralizationYong Lin, Lu Tan, Yifan Hao, Honam Wong et al.ICLR 2024 · 34 citations
- Calibration and Correctness of Language Models for CodeClaudio Spiess, David Gros, Kunal Suresh Pai, Michael Pradel et al.ICSE 2025 · 21 citations
- Self-Evolution Learning for Mixup: Enhance Data Augmentation on Few-Shot Text Classification TasksHaoqi Zheng, Qihuang Zhong, Liang Ding, Zhiliang Tian et al.EMNLP 2023 · 4 citations
Builds on6
- MixText: Linguistically-Informed Interpolation of Hidden Space for Semi-Supervised Text ClassificationJiaao Chen, Zichao Yang, Diyi YangACL 2020 · 340 citations
- On the Inference Calibration of Neural Machine TranslationShuo Wang, Zhaopeng Tu, Shuming Shi, Yang LiuACL 2020 · 66 citations
- SeqMix: Augmenting Active Sequence Labeling via Sequence MixupRongzhi Zhang, Yue Yu, Chao ZhangEMNLP 2020 · 65 citations
- Calibrated Language Model Fine-Tuning for In- and Out-of-Distribution DataLingkai Kong, Haoming Jiang, Yuchen Zhuang, Jie Lyu et al.EMNLP 2020 · 47 citations
- Calibrating Structured Output Predictors for Natural Language ProcessingAbhyuday Jagannatha, Hong YuACL 2020 · 22 citations
Related papers
- Calibrating Student Models for Emotion-related TasksMahshid Hosseini, Cornelia CarageaEMNLP 2022 · 2 citations
- Preserving Pre-trained Features Helps Calibrate Fine-tuned Language ModelsGuande He, Jianfei Chen, Jun ZhuICLR 2023 · 1 citation
- On the Limitations of Temperature Scaling for Distributions with OverlapsMuthu Chidambaram, Rong GeICLR 2024 · 11 citations
- When and How Mixup Improves CalibrationLinjun Zhang, Zhun Deng, Kenji Kawaguchi, James ZouICML 2022 · 79 citations
- A Close Look into the Calibration of Pre-trained Language ModelsYangyi Chen, Lifan Yuan, Ganqu Cui, Zhiyuan Liu et al.ACL 2023 · 12 citations
