More or Less: When and How to Build Convolutional Neural Network Ensembles
Abdul Wasay, Stratos Idreos
Abstract
Convolutional neural networks are utilized to solve increasingly more complex problems and with more data. As a result, researchers and practitioners seek to scale the representational power of such models by adding more parameters. However, increasing parameters requires additional critical resources in terms of memory and compute, leading to increased training and inference cost. Thus a consistent challenge is to obtain as high as possible accuracy within a parameter budget. As neural network designers navigate this complex landscape, they are guided by conventional wisdom that is informed from past empirical studies. We identify a critical part of this design space that is not well-understood: How to decide between the alternatives of expanding a single convolutional network model or increasing the number of networks in the form of an ensemble. We study this question in detail across various network architectures and data sets. We build an extensive experimental framework that captures numerous angles of the possible design space in terms of how a new set of parameters can be used in a model. We consider a holistic set of metrics such as training time, inference time, and memory usage. The framework provides a robust assessment by making sure it controls for the number of parameters. Contrary to conventional wisdom, we show that when we perform a holistic and robust assessment, we uncover a wide design space, where ensembles provide better accuracy, train faster, and deploy at speed comparable to single convolutional networks with the same total number of parameters.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get edc9cb91-98c9-422d-b46f-1497d31610ccCited by top-tier papers4
- Deep Ensembles Work, But Are They Necessary?Taiga Abe, Estefany Kelly Buchanan, Geoff Pleiss, Richard S. Zemel et al.NeurIPS 2022 · 101 citations
- Boosting Randomized Smoothing with Variance Reduced ClassifiersMiklós Z. Horváth, Mark Niklas Müller, Marc Fischer, Martin T. VechevICLR 2022 · 56 citations
- Cosine: A Cloud-Cost Optimized Self-Designing Key-Value Storage EngineSubarna Chatterjee, Meena Jagadeesan, Wilson Qin, Stratos IdreosVLDB 2022 · 17 citations
- Deep Ensembles for Graphs with Higher-order DependenciesSteven J. Krieg, William C. Burgis, Patrick M. Soga, Nitesh V. ChawlaICLR 2023 · 1 citation
Related papers
- Fast and Accurate Model ScalingPiotr Dollár, Mannat Singh, Ross B. GirshickCVPR 2021
- Collegial EnsemblesEtai Littwin, Ben Myara, Sima Sabah, Joshua M. Susskind et al.NeurIPS 2020 · 10 citations
- Theoretical Limitations of Ensembles in the Age of OverparameterizationNiclas Dern, John Patrick Cunningham, Geoff PleissICML 2025
- Improving Ensemble Distillation With Weight Averaging and Diversifying PerturbationGiung Nam, Hyungi Lee, Byeongho Heo, Juho LeeICML 2022 · 10 citations
- Ex Uno Pluria: Insights on Ensembling in Low Precision Number SystemsGiung Nam, Juho LeeNeurIPS 2024 · 2 citations
