Functional Ensemble Distillation
Coby Penso, Idan Achituve, Ethan Fetaya
Abstract
Bayesian models have many desirable properties, most notable is their ability to generalize from limited data and to properly estimate the uncertainty in their predictions. However, these benefits come at a steep computational cost as Bayesian inference, in most cases, is computationally intractable. One popular approach to alleviate this problem is using a Monte-Carlo estimation with an ensemble of models sampled from the posterior. However, this approach still comes at a significant computational cost, as one needs to store and run multiple models at test time. In this work, we investigate how to best distill an ensemble's predictions using an efficient model. First, we argue that current approaches that simply return distribution over predictions cannot compute important properties, such as the covariance between predictions, which can be valuable for further processing. Second, in many limited data settings, all ensemble members achieve nearly zero training loss, namely, they produce near-identical predictions on the training set which results in sub-optimal distilled models. To address both problems, we propose a novel and general distillation approach, named Functional Ensemble Distillation (FED), and we investigate how to best distill an ensemble in this setting. We find that learning the distilled model via a simple augmentation scheme in the form of mixup augmentation [42] significantly boosts the performance. We evaluated our method on several tasks and showed that it achieves superior results in both accuracy and uncertainty estimation compared to current approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 912b0b04-e3fa-48c8-8e0a-8d78e92d2e38Cited by top-tier papers3
- Fast Ensembling with Diffusion Schrödinger BridgeHyunsu Kim, Jongmin Yoon, Juho LeeICLR 2024 · 2 citations
- Efficient Credal Prediction through DecalibrationPaul Hofman, Timo Löhr, Maximilian Muschalik, Yusuf Sale et al.ICLR 2026 · 1 citation
- Ensemble Distribution Distillation via Flow MatchingJonggeon Park, Giung Nam, Hyunsu Kim, Jongmin Yoon et al.ICML 2025
Builds on10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- BatchEnsemble: an Alternative Approach to Efficient Ensemble and Lifelong LearningYeming Wen, Dustin Tran, Jimmy BaICLR 2020 · 569 citations
- What Are Bayesian Neural Network Posteriors Really Like?Pavel Izmailov, Sharad Vikram, Matthew D. Hoffman, Andrew Gordon WilsonICML 2021 · 458 citations
- Cyclical Stochastic Gradient MCMC for Bayesian Deep LearningRuqi Zhang, Chunyuan Li, Jianyi Zhang, Changyou Chen et al.ICLR 2020 · 292 citations
- Ensemble Distribution DistillationAndrey Malinin, Bruno Mlodozeniec, Mark J. F. GalesICLR 2020 · 273 citations
Related papers
- Bayesian Knowledge Distillation: A Bayesian Perspective of Distillation with Uncertainty QuantificationLuyang Fang, Yongkai Chen, Wenxuan Zhong, Ping MaICML 2024 · 10 citations
- Improving Ensemble Distillation With Weight Averaging and Diversifying PerturbationGiung Nam, Hyungi Lee, Byeongho Heo, Juho LeeICML 2022 · 10 citations
- Diversity Matters When Learning From EnsemblesGiung Nam, Jongmin Yoon, Yoonho Lee, Juho LeeNeurIPS 2021 · 50 citations
- Probabilistic Knowledge Distillation of Face EnsemblesJianqing Xu, Shen Li, Ailin Deng, Miao Xiong et al.CVPR 2023
- MixKD: Towards Efficient Distillation of Large-scale Language ModelsKevin J. Liang, Weituo Hao, Dinghan Shen, Yufan Zhou et al.ICLR 2021 · 90 citations
