Lottery Tickets on a Data Diet: Finding Initializations with Sparse Trainable Networks
Mansheej Paul, Brett W. Larsen, Surya Ganguli, Jonathan Frankle, Gintare Karolina Dziugaite
摘要
A striking observation about iterative magnitude pruning (IMP; Frankle et al. 2020) is that after just a few hundred steps of dense training the method can find a sparse sub-network that can be trained to the same accuracy as the dense network. However, the same does not hold at step 0, i.e. random initialization. In this work, we seek to understand how this early phase of pre-training leads to a good initialization for IMP both through the lens of the data distribution and the loss landscape geometry. Empirically we observe that, holding the number of pre-training iterations constant, training on a small fraction of (randomly chosen) data suffices to obtain an equally good initialization for IMP. We additionally observe that by pre-training only on"easy"training data, we can decrease the number of steps necessary to find a good initialization for IMP compared to training on the full dataset or a randomly chosen subset. Finally, we identify novel properties of the loss landscape of dense networks that are predictive of IMP performance, showing in particular that more examples being linearly mode connected in the dense network correlates well with good initializations for IMP. Combined, these results provide new insight into the role played by the early phase training in IMP.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Masks, Signs, And Learning Rate RewindingAdvait Harshal Gadhikar, Rebekka BurkholzICLR 2024 · 被引用 15 次
- Spanning Training Progress: Temporal Dual-Depth Scoring (TDDS) for Enhanced Dataset PruningXin Zhang, Jiawei Du, Yunsong Li, Weiying Xie 等CVPR 2024 · 被引用 12 次
- Evolution-aware VAriance (EVA) Coreset Selection for Medical Image ClassificationYuxin Hong, Xiao Zhang, Xin Zhang, Joey Tianyi ZhouACM MM 2024 · 被引用 4 次
- Robust Active DistillationCenk Baykal, Khoa Trinh, Fotis Iliopoulos, Gaurav Menghani 等ICLR 2023 · 被引用 3 次
- Unmasking the Lottery Ticket Hypothesis: What's Encoded in a Winning Ticket's Mask?Mansheej Paul, Feng Chen, Brett W. Larsen, Jonathan Frankle 等ICLR 2023 · 被引用 2 次
它引用的顶会 Paper12
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 被引用 806 次
- Linear Mode Connectivity and the Lottery Ticket HypothesisJonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, Michael CarbinICML 2020 · 被引用 750 次
- Rigging the Lottery: Making All Tickets WinnersUtku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro 等ICML 2020 · 被引用 723 次
- The Lottery Ticket Hypothesis for Pre-trained BERT NetworksTianlong Chen, Jonathan Frankle, Shiyu Chang, Sijia Liu 等NeurIPS 2020 · 被引用 428 次
- The Role of Permutation Invariance in Linear Mode Connectivity of Neural NetworksRahim Entezari, Hanie Sedghi, Olga Saukh, Behnam NeyshaburICLR 2022 · 被引用 301 次
相关 Paper
- Neural Network Pruning Denoises the Features and Makes Local Connectivity Emerge in Visual TasksFranco Pellegrini, Giulio BiroliICML 2022 · 被引用 10 次
- Rare Gems: Finding Lottery Tickets at InitializationKartik Sreenivasan, Jy-yong Sohn, Liu Yang, Matthew Grinde 等NeurIPS 2022 · 被引用 52 次
- When to Prune? A Policy towards Early Structural PruningMaying Shen, Pavlo Molchanov, Hongxu Yin, José M. ÁlvarezCVPR 2022 · 被引用 46 次
- No Free Prune: Information-Theoretic Barriers to Pruning at InitializationTanishq Kumar, Kevin Luo, Mark SellkeICML 2024 · 被引用 9 次
- How I Learned to Stop Worrying and Love RetrainingMax Zimmer, Christoph Spiegel, Sebastian PokuttaICLR 2023
