Deep Ensembling with No Overhead for either Training or Testing: The All-Round Blessings of Dynamic Sparsity
Shiwei Liu, Tianlong Chen, Zahra Atashgahi, Xiaohan Chen, Ghada Sokar, Elena Mocanu, Mykola Pechenizkiy, Zhangyang Wang, Decebal Constantin Mocanu
Abstract
The success of deep ensembles on improving predictive performance, uncertainty estimation, and out-of-distribution robustness has been extensively studied in the machine learning literature. Albeit the promising results, naively training multiple deep neural networks and combining their predictions at inference leads to prohibitive computational costs and memory requirements. Recently proposed efficient ensemble approaches reach the performance of the traditional deep ensembles with significantly lower costs. However, the training resources required by these approaches are still at least the same as training a single dense model. In this work, we draw a unique connection between sparse neural network training and deep ensembles, yielding a novel efficient ensemble learning framework called F reeT ickets. Instead of training multiple dense networks and averaging them, we directly train sparse subnetworks from scratch and extract diverse yet accurate subnetworks during this efficient, sparse-to-sparse training. Our framework, F reeT ickets, is defined as the ensemble of these relatively cheap sparse subnetworks. Despite being an ensemble method, F reeT ickets has even fewer parameters and training FLOPs than a single dense model. This seemingly counterintuitive outcome is due to the ultra training/inference efficiency of dynamic sparse training. F reeT ickets surpasses the dense baseline in all the following criteria: prediction accuracy, uncertainty estimation, out-of-distribution (OoD) robustness, as well as efficiency for both training and inference. Impressively, F reeT ickets outperforms the naive deep ensemble with ResNet50 on ImageNet using around only 1/5 of the training FLOPs required by the latter. We have released our source code at https://github.com/VITA-Group/FreeTickets .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b858d712-1e2b-429b-bf93-a426d283874eCited by top-tier papers23
- CARD: Classification and Regression Diffusion ModelsXizewen Han, Huangjie Zheng, Mingyuan ZhouNeurIPS 2022 · 185 citations
- Federated Dynamic Sparse Training: Computing Less, Communicating Less, Yet Learning BetterSameer Bibikar, Haris Vikalo, Zhangyang Wang, Xiaohan ChenAAAI 2022 · 133 citations
- Diffusion-Based Probabilistic Uncertainty Estimation for Active Domain AdaptationZhekai Du, Jingjing LiNeurIPS 2023 · 32 citations
- Where to Pay Attention in Sparse Training for Feature Selection?Ghada Sokar, Zahra Atashgahi, Mykola Pechenizkiy, Decebal Constantin MocanuNeurIPS 2022 · 25 citations
- Sparse Model Soups: A Recipe for Improved Pruning via Model AveragingMax Zimmer, Christoph Spiegel, Sebastian PokuttaICLR 2024 · 22 citations
Builds on27
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Pruning neural networks without any data by iteratively conserving synaptic flowHidenori Tanaka, Daniel Kunin, Daniel L. K. Yamins, Surya GanguliNeurIPS 2020 · 884 citations
- Linear Mode Connectivity and the Lottery Ticket HypothesisJonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, Michael CarbinICML 2020 · 750 citations
Related papers
- Distributionally Robust Ensemble of Lottery Tickets Towards Calibrated Sparse Network TrainingHitesh Sapkota, Dingrong Wang, Zhiqiang Tao, Qi YuNeurIPS 2023 · 5 citations
- Lottery Pools: Winning More by Interpolating Tickets without Increasing Training or Inference CostLu Yin, Shiwei Liu, Meng Fang, Tianjin Huang et al.AAAI 2023 · 14 citations
- Training independent subnetworks for robust predictionMarton Havasi, Rodolphe Jenatton, Stanislav Fort, Jeremiah Zhe Liu et al.ICLR 2021 · 235 citations
- Sparsity Winning Twice: Better Robust Generalization from More Efficient TrainingTianlong Chen, Zhenyu Zhang, Pengjun Wang, Santosh Balachandra et al.ICLR 2022 · 54 citations
- Winning the Lottery with Continuous SparsificationPedro Savarese, Hugo Silva, Michael MaireNeurIPS 2020 · 162 citations
