On the Calibration of Pre-trained Language Models using Mixup Guided by Area Under the Margin and Saliency
Seoyeon Park, Cornelia Caragea
摘要
A well-calibrated neural model produces confidence (probability outputs) closely approximated by the expected accuracy. While prior studies have shown that mixup training as a data augmentation technique can improve model calibration on image classification tasks, little is known about using mixup for model calibration on natural language understanding (NLU) tasks. In this paper, we explore mixup for model calibration on several NLU tasks and propose a novel mixup strategy for pre-trained language models that improves model calibration further. Our proposed mixup is guided by both the Area Under the Margin (AUM) statistic (Pleiss et al., 2020) and the saliency map of each sample (Simonyan et al., 2013). Moreover, we combine our mixup strategy with model miscalibration correction techniques (i.e., label smoothing and temperature scaling) and provide detailed analyses of their impact on our proposed mixup. We focus on systematically designing experiments on three NLU tasks: natural language inference, paraphrase detection, and commonsense reasoning. Our method achieves the lowest expected calibration error compared to strong baselines on both in-domain and out-of-domain test samples while maintaining competitive accuracy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Bayesian Low-rank Adaptation for Large Language ModelsAdam X. Yang, Maxime Robeyns, Xi Wang, Laurence AitchisonICLR 2024 · 被引用 111 次
- Thermometer: Towards Universal Calibration for Large Language ModelsMaohao Shen, Subhro Das, Kristjan H. Greenewald, Prasanna Sattigeri 等ICML 2024 · 被引用 38 次
- Spurious Feature Diversification Improves Out-of-distribution GeneralizationYong Lin, Lu Tan, Yifan Hao, Honam Wong 等ICLR 2024 · 被引用 34 次
- Calibration and Correctness of Language Models for CodeClaudio Spiess, David Gros, Kunal Suresh Pai, Michael Pradel 等ICSE 2025 · 被引用 21 次
- Self-Evolution Learning for Mixup: Enhance Data Augmentation on Few-Shot Text Classification TasksHaoqi Zheng, Qihuang Zhong, Liang Ding, Zhiliang Tian 等EMNLP 2023 · 被引用 4 次
它引用的顶会 Paper6
- MixText: Linguistically-Informed Interpolation of Hidden Space for Semi-Supervised Text ClassificationJiaao Chen, Zichao Yang, Diyi YangACL 2020 · 被引用 340 次
- On the Inference Calibration of Neural Machine TranslationShuo Wang, Zhaopeng Tu, Shuming Shi, Yang LiuACL 2020 · 被引用 66 次
- SeqMix: Augmenting Active Sequence Labeling via Sequence MixupRongzhi Zhang, Yue Yu, Chao ZhangEMNLP 2020 · 被引用 65 次
- Calibrated Language Model Fine-Tuning for In- and Out-of-Distribution DataLingkai Kong, Haoming Jiang, Yuchen Zhuang, Jie Lyu 等EMNLP 2020 · 被引用 47 次
- Calibrating Structured Output Predictors for Natural Language ProcessingAbhyuday Jagannatha, Hong YuACL 2020 · 被引用 22 次
相关 Paper
- Calibrating Student Models for Emotion-related TasksMahshid Hosseini, Cornelia CarageaEMNLP 2022 · 被引用 2 次
- Preserving Pre-trained Features Helps Calibrate Fine-tuned Language ModelsGuande He, Jianfei Chen, Jun ZhuICLR 2023 · 被引用 1 次
- On the Limitations of Temperature Scaling for Distributions with OverlapsMuthu Chidambaram, Rong GeICLR 2024 · 被引用 11 次
- When and How Mixup Improves CalibrationLinjun Zhang, Zhun Deng, Kenji Kawaguchi, James ZouICML 2022 · 被引用 79 次
- A Close Look into the Calibration of Pre-trained Language ModelsYangyi Chen, Lifan Yuan, Ganqu Cui, Zhiyuan Liu 等ACL 2023 · 被引用 12 次
