Sparse Winning Tickets are Data-Efficient Image Recognizers
Mukund Varma T., Xuxi Chen, Zhenyu Zhang, Tianlong Chen, Subhashini Venugopalan, Zhangyang Wang
Abstract
Improving the performance of deep networks in data-limited regimes has warranted much attention. In this work, we empirically show that “winning tickets” (small sub-networks) obtained via magnitude pruning based on the lottery ticket hypothesis [1], apart from being sparse are also effective recognizers in data-limited regimes. Based on extensive experiments, we find that in low data regimes (datasets of 50-100 examples per class), sparse winning tickets substantially outperform the original dense networks. This approach, when combined with augmentations or fine-tuning from a self-supervised backbone network, shows further improvements in performance by as much as 16% (absolute) on low sample datasets and long-tailed classification. Further, sparse winning tickets are more robust to synthetic noise and distribution shifts compared to their dense counterparts. Our analysis of winning tickets on small datasets indicates that, though sparse, the networks retain density in the initial layers and their representations are more generalizable. Code is available at https://github.com/VITA-Group/DataEfficientLTH .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Data Efficient Neural Scaling Law via Model ReusingPeihao Wang, Rameswar Panda, Zhangyang WangICML 2023 · 18 citations
- Lowering the Pre-training Tax for Gradient-based Subset Training: A Lightweight Distributed Pre-Training ToolkitYeonju Ro, Zhangyang Wang, Vijay Chidambaram, Aditya AkellaICML 2023
Builds on11
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
- Rethinking ImageNet Pre-TrainingKaiming He, Ross B. Girshick, Piotr DollárICCV 2019 · 1,188 citations
- Linear Mode Connectivity and the Lottery Ticket HypothesisJonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, Michael CarbinICML 2020 · 750 citations
- Movement Pruning: Adaptive Sparsity by Fine-TuningVictor Sanh, Thomas Wolf, Alexander M. RushNeurIPS 2020 · 656 citations
Related papers
- Dual Lottery Ticket HypothesisYue Bai, Huan Wang, Zhiqiang Tao, Kunpeng Li et al.ICLR 2022 · 49 citations
- Quarantine: Sparsity Can Uncover the Trojan Attack Trigger for FreeTianlong Chen, Zhenyu Zhang, Yihua Zhang, Shiyu Chang et al.CVPR 2022 · 13 citations
- The Lottery Ticket Hypothesis for Object RecognitionSharath Girish, Shishira R. Maiya, Kamal Gupta, Hao Chen et al.CVPR 2021
- When BERT Plays the Lottery, All Tickets Are WinningSai Prasanna, Anna Rogers, Anna RumshiskyEMNLP 2020 · 114 citations
- Lottery Pools: Winning More by Interpolating Tickets without Increasing Training or Inference CostLu Yin, Shiwei Liu, Meng Fang, Tianjin Huang et al.AAAI 2023 · 14 citations
