Rare Gems: Finding Lottery Tickets at Initialization
Kartik Sreenivasan, Jy-yong Sohn, Liu Yang, Matthew Grinde, Alliot Nagle, Hongyi Wang, Eric P. Xing, Kangwook Lee, Dimitris S. Papailiopoulos
Abstract
Large neural networks can be pruned to a small fraction of their original size, with little loss in accuracy, by following a time-consuming"train, prune, re-train"approach. Frankle&Carbin conjecture that we can avoid this by training"lottery tickets", i.e., special sparse subnetworks found at initialization, that can be trained to high accuracy. However, a subsequent line of work by Frankle et al. and Su et al. presents concrete evidence that current algorithms for finding trainable networks at initialization, fail simple baseline comparisons, e.g., against training random sparse subnetworks. Finding lottery tickets that train to better accuracy compared to simple baselines remains an open problem. In this work, we resolve this open problem by proposing Gem-Miner which finds lottery tickets at initialization that beat current baselines. Gem-Miner finds lottery tickets trainable to accuracy competitive or better than Iterative Magnitude Pruning (IMP), and does so up to faster.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext efc1c8b5-99a7-4073-8f83-a0eba9cbd223Cited by top-tier papers16
- Pruner-Zero: Evolving Symbolic Pruning Metric From Scratch for Large Language ModelsPeijie Dong, Lujun Li, Zhenheng Tang, Xiang Liu et al.ICML 2024 · 64 citations
- Fantastic Weights and How to Find Them: Where to Prune in Dynamic Sparse TrainingAleksandra Nowak, Bram Grooten, Decebal Constantin Mocanu, Jacek TaborNeurIPS 2023 · 23 citations
- How Sparse Can We Prune A Deep Network: A Fundamental Limit PerspectiveQiaozhe Zhang, Ruijie Zhang, Jun Sun, Yingzhuang LiuNeurIPS 2024 · 14 citations
- Maestro: Uncovering Low-Rank Structures via Trainable DecompositionSamuel Horváth, Stefanos Laskaridis, Shashank Rajput, Hongyi WangICML 2024 · 11 citations
- No Free Prune: Information-Theoretic Barriers to Pruning at InitializationTanishq Kumar, Kevin Luo, Mark SellkeICML 2024 · 9 citations
Builds on16
- Pruning neural networks without any data by iteratively conserving synaptic flowHidenori Tanaka, Daniel Kunin, Daniel L. K. Yamins, Surya GanguliNeurIPS 2020 · 884 citations
- Linear Mode Connectivity and the Lottery Ticket HypothesisJonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, Michael CarbinICML 2020 · 750 citations
- Picking Winning Tickets Before Training by Preserving Gradient FlowChaoqi Wang, Guodong Zhang, Roger B. GrosseICLR 2020 · 743 citations
- Comparing Rewinding and Fine-tuning in Neural Network PruningAlex Renda, Jonathan Frankle, Michael CarbinICLR 2020 · 437 citations
- The Lottery Ticket Hypothesis for Pre-trained BERT NetworksTianlong Chen, Jonathan Frankle, Shiyu Chang, Sijia Liu et al.NeurIPS 2020 · 428 citations
Related papers
- Winning Lottery Tickets in Deep Generative ModelsNeha Mukund Kalibhat, Yogesh Balaji, Soheil FeiziAAAI 2021 · 44 citations
- Efficient Lottery Ticket Finding: Less Data is MoreZhenyu Zhang, Xuxi Chen, Tianlong Chen, Zhangyang WangICML 2021 · 58 citations
- The Elastic Lottery Ticket HypothesisXiaohan Chen, Yu Cheng, Shuohang Wang, Zhe Gan et al.NeurIPS 2021 · 38 citations
- Proving the Lottery Ticket Hypothesis for Convolutional Neural NetworksArthur da Cunha, Emanuele Natale, Laurent ViennotICLR 2022 · 31 citations
- Winning the Lottery with Continuous SparsificationPedro Savarese, Hugo Silva, Michael MaireNeurIPS 2020 · 162 citations
