Diversity Matters When Learning From Ensembles
Giung Nam, Jongmin Yoon, Yoonho Lee, Juho Lee
Abstract
Deep ensembles excel in large-scale image classification tasks both in terms of prediction accuracy and calibration. Despite being simple to train, the computation and memory cost of deep ensembles limits their practicability. While some recent works propose to distill an ensemble model into a single model to reduce such costs, there is still a performance gap between the ensemble and distilled models. We propose a simple approach for reducing this gap, i.e., making the distilled performance close to the full ensemble. Our key assumption is that a distilled model should absorb as much function diversity inside the ensemble as possible. We first empirically show that the typical distillation procedure does not effectively transfer such diversity, especially for complex models that achieve near-zero training error. To fix this, we propose a perturbation strategy for distillation that reveals diversity by seeking inputs for which ensemble member outputs disagree. We empirically show that a model distilled with such perturbed samples indeed exhibits enhanced diversity, leading to improved performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4111ab0c-d6f6-468e-8cc3-82729ee77d37Cited by top-tier papers11
- Continual Variational Autoencoder via Continual Generative Knowledge DistillationFei Ye, Adrian G. BorsAAAI 2023 · 13 citations
- Improving Ensemble Distillation With Weight Averaging and Diversifying PerturbationGiung Nam, Hyungi Lee, Byeongho Heo, Juho LeeICML 2022 · 10 citations
- Fed-DFA: Federated Distillation for Heterogeneous Model Fusion Through the Adversarial LensZichen Wang, Feng Yan, Tianyi Wang, Cong Wang et al.AAAI 2025 · 9 citations
- Top-Ambiguity Samples Matter: Understanding Why Deep Ensemble Works in Selective ClassificationQiang Ding, Yixuan Cao, Ping LuoNeurIPS 2023 · 6 citations
- Functional Ensemble DistillationCoby Penso, Idan Achituve, Ethan FetayaNeurIPS 2022 · 3 citations
Builds on7
- Bayesian Deep Learning and a Probabilistic Perspective of GeneralizationAndrew Gordon Wilson, Pavel IzmailovNeurIPS 2020 · 845 citations
- On the Efficacy of Knowledge DistillationJang Hyun Cho, Bharath HariharanICCV 2019 · 741 citations
- BatchEnsemble: an Alternative Approach to Efficient Ensemble and Lifelong LearningYeming Wen, Dustin Tran, Jimmy BaICLR 2020 · 569 citations
- Pitfalls of In-Domain Uncertainty Estimation and Ensembling in Deep LearningArsenii Ashukha, Alexander Lyzhov, Dmitry Molchanov, Dmitry P. VetrovICLR 2020 · 354 citations
- Ensemble Distribution DistillationAndrey Malinin, Bruno Mlodozeniec, Mark J. F. GalesICLR 2020 · 273 citations
Related papers
- Ensemble Distribution Distillation via Flow MatchingJonggeon Park, Giung Nam, Hyunsu Kim, Jongmin Yoon et al.ICML 2025
- Repulsive Deep Ensembles are BayesianFrancesco D'Angelo, Vincent FortuinNeurIPS 2021 · 141 citations
- Joint Training of Deep Ensembles Fails Due to Learner CollusionAlan Jeffares, Tennison Liu, Jonathan Crabbé, Mihaela van der SchaarNeurIPS 2023 · 34 citations
- DICE: Diversity in Deep Ensembles via Conditional Redundancy Adversarial EstimationAlexandre Ramé, Matthieu CordICLR 2021 · 60 citations
- Regularizing Class-Wise Predictions via Self-Knowledge DistillationSukmin Yun, Jongjin Park, Kimin Lee, Jinwoo ShinCVPR 2020
