Scaling Ensemble Distribution Distillation to Many Classes with Proxy Targets
Max Ryabinin, Andrey Malinin, Mark J. F. Gales
Abstract
Ensembles of machine learning models yield improved system performance as well as robust and interpretable uncertainty estimates; however, their inference costs may often be prohibitively high. Ensemble Distribution Distillation is an approach that allows a single model to efficiently capture both the predictive performance and uncertainty estimates of an ensemble. For classification, this is achieved by training a Dirichlet distribution over the ensemble members' output distributions via the maximum likelihood criterion. Although theoretically principled, this criterion exhibits poor convergence when applied to large-scale tasks where the number of classes is very high. In our work, we analyze this effect and show that for the Dirichlet log-likelihood criterion classes with low probability induce larger gradients than high-probability classes. This forces the model to focus on the distribution of the ensemble tail-class probabilities. We propose a new training objective which minimizes the reverse KL-divergence to a Proxy-Dirichlet target derived from the ensemble. This loss resolves the gradient issues of Ensemble Distribution Distillation, as we demonstrate both theoretically and empirically on the ImageNet and WMT17 En-De datasets containing 1000 and 40,000 classes, respectively. * Equal contribution. 2 Data and Knowledge Uncertainty are also known as Aleatoric and Epistemic uncertainty. Preprint. Under review.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 04f209fb-c35f-4faa-b840-5e5ac17dc175Cited by top-tier papers9
- Window-Based Early-Exit Cascades for Uncertainty Estimation: When Deep Ensembles are More Efficient than Single ModelsGuoxuan Xia, Christos-Savvas BouganisICCV 2023 · 17 citations
- Improving Ensemble Distillation With Weight Averaging and Diversifying PerturbationGiung Nam, Hyungi Lee, Byeongho Heo, Juho LeeICML 2022 · 10 citations
- Uncertainty Measures in Neural Belief Tracking and the Effects on Dialogue Policy PerformanceCarel van Niekerk, Andrey Malinin, Christian Geishauser, Michael Heck et al.EMNLP 2021 · 8 citations
- Evidential Knowledge DistillationLiangyu Xiang, Junyu Gao, Changsheng XuICCV 2025 · 6 citations
- Fast Ensembling with Diffusion Schrödinger BridgeHyunsu Kim, Jongmin Yoon, Juho LeeICLR 2024 · 2 citations
Builds on4
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath et al.ICCV 2021 · 2,294 citations
- Pitfalls of In-Domain Uncertainty Estimation and Ensembling in Deep LearningArsenii Ashukha, Alexander Lyzhov, Dmitry Molchanov, Dmitry P. VetrovICLR 2020 · 354 citations
- Ensemble Distribution DistillationAndrey Malinin, Bruno Mlodozeniec, Mark J. F. GalesICLR 2020 · 273 citations
- Natural Adversarial ExamplesDan Hendrycks, Kevin Zhao, Steven Basart, Jacob Steinhardt et al.CVPR 2021
Related papers
- Gradient Reweighting: Towards Imbalanced Class-Incremental LearningJiangpeng HeCVPR 2024
- Diversity Matters When Learning From EnsemblesGiung Nam, Jongmin Yoon, Yoonho Lee, Juho LeeNeurIPS 2021 · 50 citations
- Knowledge Distillation of Uncertainty using Deep Latent Factor ModelSehyun Park, Jongjin Lee, Yunseop Shin, Ilsang Ohn et al.NeurIPS 2025 · 2 citations
- Ensemble Distribution Distillation via Flow MatchingJonggeon Park, Giung Nam, Hyunsu Kim, Jongmin Yoon et al.ICML 2025
- A Consistent and Differentiable Lp Canonical Calibration Error EstimatorTeodora Popordanoska, Raphael Sayer, Matthew B. BlaschkoNeurIPS 2022 · 58 citations
