When are ensembles really effective?
Ryan Theisen, Hyunsuk Kim, Yaoqing Yang, Liam Hodgkinson, Michael W. Mahoney
Abstract
Ensembling has a long history in statistical data analysis, with many impactful applications. However, in many modern machine learning settings, the benefits of ensembling are less ubiquitous and less obvious. We study, both theoretically and empirically, the fundamental question of when ensembling yields significant performance improvements in classification tasks. Theoretically, we prove new results relating the ensemble improvement rate (a measure of how much ensembling decreases the error rate versus a single model, on a relative scale) to the disagreement-error ratio. We show that ensembling improves performance significantly whenever the disagreement rate is large relative to the average error rate; and that, conversely, one classifier is often enough whenever the disagreement rate is low relative to the average error rate. On the way to proving these results, we derive, under a mild condition called competence, improved upper and lower bounds on the average test error rate of the majority vote classifier. To complement this theory, we study ensembling empirically in a variety of settings, verifying the predictions made by our theory, and identifying practical scenarios where ensembling does and does not result in large performance improvements. Perhaps most notably, we demonstrate a distinct difference in behavior between interpolating models (popular in current practice) and non-interpolating models (such as tree-based methods, where ensembling is popular), demonstrating that ensembling helps considerably more in the latter case than in the former.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 75117164-67e8-4db7-8549-ae207d382c8bCited by top-tier papers5
- Single-Teacher View Augmentation: Boosting Knowledge Distillation via Angular DiversitySeonghoon Yu, Dongjun Nam, Dina Katabi, Jeany SonNeurIPS 2025 · 4 citations
- Unveiling Multi-regime Patterns in SciML: Distinct Failure Modes and Regime-specific OptimizationYuxin Wang, Yuanzhe Hu, Xiaokun Zhong, Xiaopeng Wang et al.ICML 2026
- No Free Lunch from Random Feature Ensembles: Scaling Laws and Near-Optimality ConditionsBenjamin S. Ruben, William Lingxiao Tong, Hamza Tahir Chaudhry, Cengiz PehlevanICML 2025
- Agree to Disagree: Demystifying Homogeneous Deep Ensembles through Distributional EquivalenceYipei Wang, Xiaoqian WangICLR 2025
- ILIAS: Instance-Level Image retrieval At ScaleGiorgos Kordopatis-Zilos, Vladan Stojnic, Anna Manko, Pavel Suma et al.CVPR 2025
Builds on7
- What Are Bayesian Neural Network Posteriors Really Like?Pavel Izmailov, Sharad Vikram, Matthew D. Hoffman, Andrew Gordon WilsonICML 2021 · 458 citations
- Pitfalls of In-Domain Uncertainty Estimation and Ensembling in Deep LearningArsenii Ashukha, Alexander Lyzhov, Dmitry Molchanov, Dmitry P. VetrovICLR 2020 · 354 citations
- The MultiBERTs: BERT Reproductions for Robustness AnalysisThibault Sellam, Steve Yadlowsky, Ian Tenney, Jason Wei et al.ICLR 2022 · 106 citations
- Deep Ensembles Work, But Are They Necessary?Taiga Abe, Estefany Kelly Buchanan, Geoff Pleiss, Richard S. Zemel et al.NeurIPS 2022 · 101 citations
- Dangers of Bayesian Model Averaging under Covariate ShiftPavel Izmailov, Patrick Nicholson, Sanae Lotfi, Andrew Gordon WilsonNeurIPS 2021 · 51 citations
Related papers
- How many classifiers do we need?Hyunsuk Kim, Liam Hodgkinson, Ryan Theisen, Michael W. MahoneyNeurIPS 2024
- Rethinking Fano's Inequality in Ensemble LearningTerufumi Morishita, Gaku Morio, Shota Horiguchi, Hiroaki Ozaki et al.ICML 2022 · 4 citations
- Joint Training of Deep Ensembles Fails Due to Learner CollusionAlan Jeffares, Tennison Liu, Jonathan Crabbé, Mihaela van der SchaarNeurIPS 2023 · 34 citations
- Top-Ambiguity Samples Matter: Understanding Why Deep Ensemble Works in Selective ClassificationQiang Ding, Yixuan Cao, Ping LuoNeurIPS 2023 · 6 citations
- Subsampled Ensemble Can Improve Generalization Tail ExponentiallyHuajie Qian, Donghao Ying, Henry Lam, Wotao YinNeurIPS 2025 · 3 citations
