Calibrating Zero-shot Cross-lingual (Un-)structured Predictions
Zhengping Jiang, Anqi Liu, Benjamin Van Durme
Abstract
We investigate model calibration in the setting of zero-shot cross-lingual transfer with large-scale pre-trained language models. The level of model calibration is an important metric for evaluating the trustworthiness of predictive models. There exists an essential need for model calibration when natural language models are deployed in critical tasks. We study different post-training calibration methods in structured and unstructured prediction tasks. We find that models trained with data from the source language become less calibrated when applied to the target language and that calibration errors increase with intrinsic task difficulty and relative sparsity of training data. Moreover, we observe a potential connection between the level of calibration error and an earlier proposed measure of the distance from English to other languages. Finally, our comparison demonstrates that among other methods Temperature Scaling (TS) generalizes well to distant languages, but TS fails to calibrate more complex confidence estimation in structured predictions compared to more expressive alternatives like Gaussian Process Calibration.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2b468da8-8d8e-4c5d-8fa7-2b77379e5d76Cited by top-tier papers1
Ask how each one uses itBuilds on18
- Calibrate Before Use: Improving Few-shot Performance of Language ModelsZihao Zhao, Eric Wallace, Shi Feng, Dan Klein et al.ICML 2021 · 1,843 citations
- Your classifier is secretly an energy based model and you should treat it like oneWill Grathwohl, Kuan-Chieh Wang, Jörn-Henrik Jacobsen, David Duvenaud et al.ICLR 2020 · 643 citations
- Revisiting the Calibration of Modern Neural NetworksMatthias Minderer, Josip Djolonga, Rob Romijnders, Frances Hubis et al.NeurIPS 2021 · 633 citations
- Cross-Lingual Ability of Multilingual BERT: An Empirical StudyKarthikeyan K, Zihan Wang, Stephen Mayhew, Dan RothICLR 2020 · 378 citations
- A Joint Neural Model for Information Extraction with Global FeaturesYing Lin, Heng Ji, Fei Huang, Lingfei WuACL 2020 · 376 citations
Related papers
- An Empirical Study Into What Matters for Calibrating Vision-Language ModelsWeijie Tu, Weijian Deng, Dylan Campbell, Stephen Gould et al.ICML 2024 · 18 citations
- Preserving Pre-trained Features Helps Calibrate Fine-tuned Language ModelsGuande He, Jianfei Chen, Jun ZhuICLR 2023 · 1 citation
- Transferable Post-hoc Calibration on Pretrained Transformers in Noisy Text ClassificationJun Zhang, Wen Yao, Xiaoqian Chen, Ling FengAAAI 2023 · 5 citations
- A Close Look into the Calibration of Pre-trained Language ModelsYangyi Chen, Lifan Yuan, Ganqu Cui, Zhiyuan Liu et al.ACL 2023 · 12 citations
- Proximity-Informed Calibration for Deep Neural NetworksMiao Xiong, Ailin Deng, Pang Wei Koh, Jiaying Wu et al.NeurIPS 2023 · 30 citations
