Lottery Pools: Winning More by Interpolating Tickets without Increasing Training or Inference Cost
Lu Yin, Shiwei Liu, Meng Fang, Tianjin Huang, Vlado Menkovski, Mykola Pechenizkiy
摘要
Lottery tickets (LTs) is able to discover accurate and sparse subnetworks that could be trained in isolation to match the performance of dense networks. Ensemble, in parallel, is one of the oldest time-proven tricks in machine learning to improve performance by combining the output of multiple independent models. However, the benefits of ensemble in the context of LTs will be diluted since ensemble does not directly lead to stronger sparse subnetworks, but leverages their predictions for a better decision. In this work, we first observe that directly averaging the weights of the adjacent learned subnetworks significantly boosts the performance of LTs. Encouraged by this observation, we further propose an alternative way to perform an "ensemble'' over the subnetworks identified by iterative magnitude pruning via a simple interpolating strategy. We call our method Lottery Pools. In contrast to the naive ensemble which brings no performance gains to each single subnetwork, Lottery Pools yields much stronger sparse subnetworks than the original LTs without requiring any extra training or inference cost. Across various modern architectures on CIFAR-10/100 and ImageNet, we show that our method achieves significant performance gains in both, in-distribution and out-of-distribution scenarios. Impressively, evaluated with VGG-16 and ResNet-18, the produced sparse subnetworks outperform the original LTs by up to 1.88% on CIFAR-100 and 2.36% on CIFAR-100-C; the resulting dense network surpasses the pre-trained dense-model up to 2.22% on CIFAR-100 and 2.38% on CIFAR-100-C. Our source code can be found at https://github.com/luuyin/Lottery-pools.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Dynamic Sparse No Training: Training-Free Fine-tuning for Sparse LLMsYuxin Zhang, Lirui Zhao, Mingbao Lin, Yunyun Sun 等ICLR 2024 · 被引用 78 次
- The Emergence of Essential Sparsity in Large Pre-trained Models: The Weights that MatterAjay Jaiswal, Shiwei Liu, Tianlong Chen, Zhangyang WangNeurIPS 2023 · 被引用 57 次
- Instant Soup: Cheap Pruning Ensembles in A Single Pass Can Draw Lottery Tickets from Large ModelsAjay Kumar Jaiswal, Shiwei Liu, Tianlong Chen, Ying Ding 等ICML 2023 · 被引用 26 次
- Sparse Model Soups: A Recipe for Improved Pruning via Model AveragingMax Zimmer, Christoph Spiegel, Sebastian PokuttaICLR 2024 · 被引用 22 次
- Sparse Weight Averaging with Multiple Particles for Iterative Magnitude PruningMoonseok Choi, Hyungi Lee, Giung Nam, Juho LeeICLR 2024 · 被引用 4 次
它引用的顶会 Paper16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs 等ICML 2022 · 被引用 1,464 次
- Linear Mode Connectivity and the Lottery Ticket HypothesisJonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, Michael CarbinICML 2020 · 被引用 750 次
相关 Paper
- Coarsening the Granularity: Towards Structurally Sparse Lottery TicketsTianlong Chen, Xuxi Chen, Xiaolong Ma, Yanzhi Wang 等ICML 2022 · 被引用 42 次
- Dual Lottery Ticket HypothesisYue Bai, Huan Wang, Zhiqiang Tao, Kunpeng Li 等ICLR 2022 · 被引用 49 次
- The Elastic Lottery Ticket HypothesisXiaohan Chen, Yu Cheng, Shuohang Wang, Zhe Gan 等NeurIPS 2021 · 被引用 38 次
- Winning the Lottery with Continuous SparsificationPedro Savarese, Hugo Silva, Michael MaireNeurIPS 2020 · 被引用 162 次
- Deep Ensembling with No Overhead for either Training or Testing: The All-Round Blessings of Dynamic SparsityShiwei Liu, Tianlong Chen, Zahra Atashgahi, Xiaohan Chen 等ICLR 2022 · 被引用 62 次
