Theoretical Limitations of Ensembles in the Age of Overparameterization
Niclas Dern, John Patrick Cunningham, Geoff Pleiss
摘要
Classic ensembles generalize better than any single component model. In contrast, recent empirical studies find that modern ensembles of (overparameterized) neural networks may not provide any inherent generalization advantage over single but larger neural networks. This paper clarifies how modern overparameterized ensembles differ from their classic underparameterized counterparts, using ensembles of random feature (RF) regressors as a basis for developing theory. In contrast to the underparameterized regime, where ensembling typically induces regularization and increases generalization, we prove with minimal assumptions that infinite ensembles of overparameterized RF regressors become pointwise equivalent to (single) infinite-width RF regressors, and finite width ensembles rapidly converge to single models with the same parameter budget. These results, which are exact for ridgeless models and approximate for small ridge penalties, imply that overparameterized ensembles and single large models exhibit nearly identical generalization. We further characterize the predictive variance amongst ensemble members, demonstrating that it quantifies the expected effects of increasing capacity rather than capturing any conventional notion of uncertainty. Our results challenge common assumptions about the advantages of ensembles in overparameterized settings, prompting a reconsideration of how well intuitions from underparameterized ensembles transfer to deep ensembles and the overparameterized regime.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Variational Deep Learning via Implicit RegularizationJonathan Wenger, Beau Coker, Juraj Marusic, John Patrick CunninghamICLR 2026 · 被引用 1 次
- No Free Lunch from Random Feature Ensembles: Scaling Laws and Near-Optimality ConditionsBenjamin S. Ruben, William Lingxiao Tong, Hamza Tahir Chaudhry, Cengiz PehlevanICML 2025
它引用的顶会 Paper10
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon 等NeurIPS 2020 · 被引用 989 次
- Finite Versus Infinite Neural Networks: an Empirical StudyJaehoon Lee, Samuel S. Schoenholz, Jeffrey Pennington, Ben Adlam 等NeurIPS 2020 · 被引用 245 次
- Understanding Double Descent Requires A Fine-Grained Bias-Variance DecompositionBen Adlam, Jeffrey PenningtonNeurIPS 2020 · 被引用 111 次
- Deep Ensembles Work, But Are They Necessary?Taiga Abe, Estefany Kelly Buchanan, Geoff Pleiss, Richard S. Zemel 等NeurIPS 2022 · 被引用 101 次
- Implicit Regularization of Random Feature ModelsArthur Jacot, Berfin Simsek, Francesco Spadaro, Clément Hongler 等ICML 2020 · 被引用 83 次
相关 Paper
- More is Better: when Infinite Overparameterization is Optimal and Overfitting is ObligatoryJames B. Simon, Dhruva Karkada, Nikhil Ghosh, Mikhail BelkinICLR 2024 · 被引用 7 次
- High-Dimensional Analysis for Generalized Nonlinear Regression: From Asymptotics to AlgorithmJian Li, Yong Liu, Weiping WangAAAI 2024 · 被引用 4 次
- Disentangling the Predictive Variance of Deep Ensembles through the Neural Tangent KernelSeijin Kobayashi, Pau Vilimelis Aceituno, Johannes von OswaldNeurIPS 2022 · 被引用 4 次
- Double Trouble in Double Descent: Bias and Variance(s) in the Lazy RegimeStéphane d'Ascoli, Maria Refinetti, Giulio Biroli, Florent KrzakalaICML 2020 · 被引用 163 次
- The Onset of Variance-Limited Behavior for Networks in the Lazy and Rich RegimesAlexander B. Atanasov, Blake Bordelon, Sabarish Sainathan, Cengiz PehlevanICLR 2023 · 被引用 4 次
