Theoretical Limitations of Ensembles in the Age of Overparameterization
Niclas Dern, John Patrick Cunningham, Geoff Pleiss
Abstract
Classic ensembles generalize better than any single component model. In contrast, recent empirical studies find that modern ensembles of (overparameterized) neural networks may not provide any inherent generalization advantage over single but larger neural networks. This paper clarifies how modern overparameterized ensembles differ from their classic underparameterized counterparts, using ensembles of random feature (RF) regressors as a basis for developing theory. In contrast to the underparameterized regime, where ensembling typically induces regularization and increases generalization, we prove with minimal assumptions that infinite ensembles of overparameterized RF regressors become pointwise equivalent to (single) infinite-width RF regressors, and finite width ensembles rapidly converge to single models with the same parameter budget. These results, which are exact for ridgeless models and approximate for small ridge penalties, imply that overparameterized ensembles and single large models exhibit nearly identical generalization. We further characterize the predictive variance amongst ensemble members, demonstrating that it quantifies the expected effects of increasing capacity rather than capturing any conventional notion of uncertainty. Our results challenge common assumptions about the advantages of ensembles in overparameterized settings, prompting a reconsideration of how well intuitions from underparameterized ensembles transfer to deep ensembles and the overparameterized regime.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cf977cd5-3501-4505-9455-8f7bffcb31f3Cited by top-tier papers2
- Variational Deep Learning via Implicit RegularizationJonathan Wenger, Beau Coker, Juraj Marusic, John Patrick CunninghamICLR 2026 · 1 citation
- No Free Lunch from Random Feature Ensembles: Scaling Laws and Near-Optimality ConditionsBenjamin S. Ruben, William Lingxiao Tong, Hamza Tahir Chaudhry, Cengiz PehlevanICML 2025
Builds on10
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon et al.NeurIPS 2020 · 989 citations
- Finite Versus Infinite Neural Networks: an Empirical StudyJaehoon Lee, Samuel S. Schoenholz, Jeffrey Pennington, Ben Adlam et al.NeurIPS 2020 · 245 citations
- Understanding Double Descent Requires A Fine-Grained Bias-Variance DecompositionBen Adlam, Jeffrey PenningtonNeurIPS 2020 · 111 citations
- Deep Ensembles Work, But Are They Necessary?Taiga Abe, Estefany Kelly Buchanan, Geoff Pleiss, Richard S. Zemel et al.NeurIPS 2022 · 101 citations
- Implicit Regularization of Random Feature ModelsArthur Jacot, Berfin Simsek, Francesco Spadaro, Clément Hongler et al.ICML 2020 · 83 citations
Related papers
- More is Better: when Infinite Overparameterization is Optimal and Overfitting is ObligatoryJames B. Simon, Dhruva Karkada, Nikhil Ghosh, Mikhail BelkinICLR 2024 · 7 citations
- High-Dimensional Analysis for Generalized Nonlinear Regression: From Asymptotics to AlgorithmJian Li, Yong Liu, Weiping WangAAAI 2024 · 4 citations
- Disentangling the Predictive Variance of Deep Ensembles through the Neural Tangent KernelSeijin Kobayashi, Pau Vilimelis Aceituno, Johannes von OswaldNeurIPS 2022 · 4 citations
- Double Trouble in Double Descent: Bias and Variance(s) in the Lazy RegimeStéphane d'Ascoli, Maria Refinetti, Giulio Biroli, Florent KrzakalaICML 2020 · 163 citations
- The Onset of Variance-Limited Behavior for Networks in the Lazy and Rich RegimesAlexander B. Atanasov, Blake Bordelon, Sabarish Sainathan, Cengiz PehlevanICLR 2023 · 4 citations
