Improving Ensemble Distillation With Weight Averaging and Diversifying Perturbation
Giung Nam, Hyungi Lee, Byeongho Heo, Juho Lee
Abstract
Ensembles of deep neural networks have demonstrated superior performance, but their heavy computational cost hinders applying them for resource-limited environments. It motivates distilling knowledge from the ensemble teacher into a smaller student network, and there are two important design choices for this ensemble distillation: 1) how to construct the student network, and 2) what data should be shown during training. In this paper, we propose a weight averaging technique where a student with multiple subnetworks is trained to absorb the functional diversity of ensemble teachers, but then those subnetworks are properly averaged for inference, giving a single student network with no additional inference cost. We also propose a perturbation strategy that seeks inputs from which the diversities of teachers can be better transferred to the student. Combining these two, our method significantly improves upon previous methods on various image classification tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 03e9e439-d72d-4f95-be18-948c9b20617cCited by top-tier papers3
- Knowledge Distillation of Uncertainty using Deep Latent Factor ModelSehyun Park, Jongjin Lee, Yunseop Shin, Ilsang Ohn et al.NeurIPS 2025 · 2 citations
- Ensemble Distribution Distillation via Flow MatchingJonggeon Park, Giung Nam, Hyunsu Kim, Jongmin Yoon et al.ICML 2025
- Accelerating Dataset Distillation via Model AugmentationLei Zhang, Jie Zhang, Bowen Lei, Subhabrata Mukherjee et al.CVPR 2023
Builds on14
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine et al.AAAI 2020 · 1,361 citations
- BatchEnsemble: an Alternative Approach to Efficient Ensemble and Lifelong LearningYeming Wen, Dustin Tran, Jimmy BaICLR 2020 · 569 citations
- Pitfalls of In-Domain Uncertainty Estimation and Ensembling in Deep LearningArsenii Ashukha, Alexander Lyzhov, Dmitry Molchanov, Dmitry P. VetrovICLR 2020 · 354 citations
- Ensemble Distribution DistillationAndrey Malinin, Bruno Mlodozeniec, Mark J. F. GalesICLR 2020 · 273 citations
- Hyperparameter Ensembles for Robustness and Uncertainty QuantificationFlorian Wenzel, Jasper Snoek, Dustin Tran, Rodolphe JenattonNeurIPS 2020 · 263 citations
Related papers
- Diversity Matters When Learning From EnsemblesGiung Nam, Jongmin Yoon, Yoonho Lee, Juho LeeNeurIPS 2021 · 50 citations
- Agree to Disagree: Adaptive Ensemble Knowledge Distillation in Gradient SpaceShangchen Du, Shan You, Xiaojie Li, Jianlong Wu et al.NeurIPS 2020 · 144 citations
- Customizing Student Networks From Heterogeneous Teachers via Adaptive Knowledge AmalgamationChengchao Shen, Mengqi Xue, Xinchao Wang, Jie Song et al.ICCV 2019 · 63 citations
- Online Knowledge Distillation with Diverse PeersDefang Chen, Jian-Ping Mei, Can Wang, Yan Feng et al.AAAI 2020 · 354 citations
- Towards Oracle Knowledge Distillation with Neural Architecture SearchMinsoo Kang, Jonghwan Mun, Bohyung HanAAAI 2020 · 48 citations
