Scaling Ensemble Distribution Distillation to Many Classes with Proxy Targets
Max Ryabinin, Andrey Malinin, Mark J. F. Gales
摘要
Ensembles of machine learning models yield improved system performance as well as robust and interpretable uncertainty estimates; however, their inference costs may often be prohibitively high. Ensemble Distribution Distillation is an approach that allows a single model to efficiently capture both the predictive performance and uncertainty estimates of an ensemble. For classification, this is achieved by training a Dirichlet distribution over the ensemble members' output distributions via the maximum likelihood criterion. Although theoretically principled, this criterion exhibits poor convergence when applied to large-scale tasks where the number of classes is very high. In our work, we analyze this effect and show that for the Dirichlet log-likelihood criterion classes with low probability induce larger gradients than high-probability classes. This forces the model to focus on the distribution of the ensemble tail-class probabilities. We propose a new training objective which minimizes the reverse KL-divergence to a Proxy-Dirichlet target derived from the ensemble. This loss resolves the gradient issues of Ensemble Distribution Distillation, as we demonstrate both theoretically and empirically on the ImageNet and WMT17 En-De datasets containing 1000 and 40,000 classes, respectively. * Equal contribution. 2 Data and Knowledge Uncertainty are also known as Aleatoric and Epistemic uncertainty. Preprint. Under review.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Window-Based Early-Exit Cascades for Uncertainty Estimation: When Deep Ensembles are More Efficient than Single ModelsGuoxuan Xia, Christos-Savvas BouganisICCV 2023 · 被引用 17 次
- Improving Ensemble Distillation With Weight Averaging and Diversifying PerturbationGiung Nam, Hyungi Lee, Byeongho Heo, Juho LeeICML 2022 · 被引用 10 次
- Uncertainty Measures in Neural Belief Tracking and the Effects on Dialogue Policy PerformanceCarel van Niekerk, Andrey Malinin, Christian Geishauser, Michael Heck 等EMNLP 2021 · 被引用 8 次
- Evidential Knowledge DistillationLiangyu Xiang, Junyu Gao, Changsheng XuICCV 2025 · 被引用 6 次
- Fast Ensembling with Diffusion Schrödinger BridgeHyunsu Kim, Jongmin Yoon, Juho LeeICLR 2024 · 被引用 2 次
它引用的顶会 Paper4
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath 等ICCV 2021 · 被引用 2,294 次
- Pitfalls of In-Domain Uncertainty Estimation and Ensembling in Deep LearningArsenii Ashukha, Alexander Lyzhov, Dmitry Molchanov, Dmitry P. VetrovICLR 2020 · 被引用 354 次
- Ensemble Distribution DistillationAndrey Malinin, Bruno Mlodozeniec, Mark J. F. GalesICLR 2020 · 被引用 273 次
- Natural Adversarial ExamplesDan Hendrycks, Kevin Zhao, Steven Basart, Jacob Steinhardt 等CVPR 2021
相关 Paper
- Gradient Reweighting: Towards Imbalanced Class-Incremental LearningJiangpeng HeCVPR 2024
- Diversity Matters When Learning From EnsemblesGiung Nam, Jongmin Yoon, Yoonho Lee, Juho LeeNeurIPS 2021 · 被引用 50 次
- Knowledge Distillation of Uncertainty using Deep Latent Factor ModelSehyun Park, Jongjin Lee, Yunseop Shin, Ilsang Ohn 等NeurIPS 2025 · 被引用 2 次
- Ensemble Distribution Distillation via Flow MatchingJonggeon Park, Giung Nam, Hyunsu Kim, Jongmin Yoon 等ICML 2025
- A Consistent and Differentiable Lp Canonical Calibration Error EstimatorTeodora Popordanoska, Raphael Sayer, Matthew B. BlaschkoNeurIPS 2022 · 被引用 58 次
