Ensemble Distillation for Structured Prediction: Calibrated, Accurate, Fast - Choose Three
Steven Reich, David Mueller, Nicholas Andrews
Abstract
Modern neural networks do not always produce well-calibrated predictions, even when trained with a proper scoring function such as cross-entropy. In classification settings, simple methods such as isotonic regression or temperature scaling may be used in conjunction with a held-out dataset to calibrate model outputs. However, extending these methods to structured prediction is not always straightforward or effective; furthermore, a held-out calibration set may not always be available. In this paper, we study ensemble distillation as a general framework for producing wellcalibrated structured prediction models while avoiding the prohibitive inference-time cost of ensembles. We validate this framework on two tasks: named-entity recognition and machine translation. We find that, across both tasks, ensemble distillation produces models which retain much of, and occasionally improve upon, the performance and calibration benefits of ensembles, while only requiring a single model during test-time.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7ede641c-d841-4831-985a-41dbbc9eb60bCited by top-tier papers3
- Model ensemble instead of prompt fusion: a sample-specific knowledge transfer method for few-shot prompt tuningXiangyu Peng, Chen Xing, Prafulla Kumar Choubey, Chien-Sheng Wu et al.ICLR 2023 · 5 citations
- Calibrating Zero-shot Cross-lingual (Un-)structured PredictionsZhengping Jiang, Anqi Liu, Benjamin Van DurmeEMNLP 2022 · 4 citations
- Are Data Augmentation Methods in Named Entity Recognition Applicable for Uncertainty Estimation?Wataru Hashimoto, Hidetaka Kamigaito, Taro WatanabeEMNLP 2024 · 1 citation
Builds on2
Related papers
- Rethinking Data Distillation: Do Not Overlook CalibrationDongyao Zhu, Yanbo Fang, Bowen Lei, Yiqun Xie et al.ICCV 2023 · 19 citations
- Calibrating Structured Output Predictors for Natural Language ProcessingAbhyuday Jagannatha, Hong YuACL 2020 · 22 citations
- Diversity Matters When Learning From EnsemblesGiung Nam, Jongmin Yoon, Yoonho Lee, Juho LeeNeurIPS 2021 · 50 citations
- Set Learning for Accurate and Calibrated ModelsLukas Muttenthaler, Robert A. Vandermeulen, Qiuyi Zhang, Thomas Unterthiner et al.ICLR 2024 · 4 citations
- Probability Calibration for Knowledge Graph Embedding ModelsPedro Tabacof, Luca CostabelloICLR 2020 · 49 citations
