Learning Compact Representations of Neural Networks using DiscriminAtive Masking (DAM)
Jie Bu, Arka Daw, M. Maruf, Anuj Karpatne
Abstract
A central goal in deep learning is to learn compact representations of features at every layer of a neural network, which is useful for both unsupervised representation learning and structured network pruning. While there is a growing body of work in structured pruning, current state-of-the-art methods suffer from two key limitations: (i) instability during training, and (ii) need for an additional step of fine-tuning, which is resource-intensive. At the core of these limitations is the lack of a systematic approach that jointly prunes and refines weights during training in a single stage, and does not require any fine-tuning upon convergence to achieve state-of-the-art performance. We present a novel single-stage structured pruning method termed DiscriminAtive Masking (DAM). The key intuition behind DAM is to discriminatively prefer some of the neurons to be refined during the training process, while gradually masking out other neurons. We show that our proposed DAM approach has remarkably good performance over a diverse range of applications in representation learning and structured pruning, including dimensionality reduction, recommendation system, graph representation learning, and structured pruning for image classification. We also theoretically show that the learning objective of DAM is directly related to minimizing the L 0 norm of the masking layer. All of our codes and datasets are available https://github.com/jayroxis/dam-pytorch .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4fd0f4cc-3bd2-42b5-9756-7ec9c271b947Builds on11
- Implicit Neural Representations with Periodic Activation FunctionsVincent Sitzmann, Julien N. P. Martel, Alexander W. Bergman, David B. Lindell et al.NeurIPS 2020 · 4,008 citations
- Linear Mode Connectivity and the Lottery Ticket HypothesisJonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, Michael CarbinICML 2020 · 750 citations
- Rigging the Lottery: Making All Tickets WinnersUtku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro et al.ICML 2020 · 723 citations
- AdaBelief Optimizer: Adapting Stepsizes by the Belief in Observed GradientsJuntang Zhuang, Tommy Tang, Yifan Ding, Sekhar Tatikonda et al.NeurIPS 2020 · 697 citations
- Comparing Rewinding and Fine-tuning in Neural Network PruningAlex Renda, Jonathan Frankle, Michael CarbinICLR 2020 · 437 citations
Related papers
- DARB: A Density-Adaptive Regular-Block Pruning for Deep Neural NetworksAo Ren, Tao Zhang, Yuhao Wang, Sheng Lin et al.AAAI 2020 · 11 citations
- S2HPruner: Soft-to-Hard Distillation Bridges the Discretization Gap in PruningWeihao Lin, Shengji Tang, Chong Yu, Peng Ye et al.NeurIPS 2024 · 2 citations
- Structured Compression by Weight Encryption for Unstructured Pruning and QuantizationSe Jung Kwon, Dongsoo Lee, Byeongwook Kim, Parichay Kapoor et al.CVPR 2020
- Only Train Once: A One-Shot Neural Network Training And Pruning FrameworkTianyi Chen, Bo Ji, Tianyu Ding, Biyi Fang et al.NeurIPS 2021 · 135 citations
- Effective Sparsification of Neural Networks With Global Sparsity ConstraintXiao Zhou, Weizhong Zhang, Hang Xu, Tong ZhangCVPR 2021
