Signing the Supermask: Keep, Hide, Invert
Nils Koster, Oliver Grothe, Achim Rettinger
Abstract
The exponential growth in numbers of parameters of neural networks over the past years has been accompanied by an increase in performance across several fields. However, due to their sheer size, the networks not only became difficult to interpret but also problematic to train and use in real-world applications, since hardware requirements increased accordingly. Tackling both issues, we present a novel approach that either drops a neural network's initial weights or inverts their respective sign. Put simply, a network is trained by weight selection and inversion without changing their absolute values. Our contribution extends previous work on masking by additionally sign-inverting the initial weights and follows the findings of the Lottery Ticket Hypothesis. Through this extension and adaptations of initialization methods, we achieve a pruning rate of up to 99%, while still matching or exceeding the performance of various baseline and previous models. Our approach has two main advantages. First, and most notable, signed Supermask models drastically simplify a model's structure, while still performing well on given tasks. Second, by reducing the neural network to its very foundation, we gain insights into which weights matter for performance. The code is available on GitHub.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Multicoated Supermasks Enhance Hidden NetworksYasuyuki Okoshi, Ángel López García-Arias, Kazutoshi Hirose, Kota Ando et al.ICML 2022 · 9 citations
- Masked Random Noise for Communication-Efficient Federated LearningShiwei Li, Yingyi Cheng, Haozhao Wang, Xing Tang et al.ACM MM 2024 · 7 citations
- Find A Winning Sign: Sign Is All We Need to Win the LotteryJunghun Oh, Sungyong Baik, Kyoung Mu LeeICLR 2025
- Policy Search, Retrieval, and Composition via Task Similarity in Collaborative Agentic SystemsSaptarshi Nath, Christos Peridis, Eseoghene Benjamin, Xinran Liu et al.AAAI 2026
Builds on5
- Rigging the Lottery: Making All Tickets WinnersUtku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro et al.ICML 2020 · 723 citations
- Progressive Skeletonization: Trimming more fat from a network at initializationPau de Jorge, Amartya Sanyal, Harkirat S. Behl, Philip H. S. Torr et al.ICLR 2021 · 110 citations
- Pruning Randomly Initialized Neural Networks with Iterative RandomizationDaiki Chijiwa, Shin'ya Yamaguchi, Yasutoshi Ida, Kenji Umakoshi et al.NeurIPS 2021 · 31 citations
- Multi-Prize Lottery Ticket Hypothesis: Finding Accurate Binary Neural Networks by Pruning A Randomly Weighted NetworkJames Diffenderfer, Bhavya KailkhuraICLR 2021 · 12 citations
- What's Hidden in a Randomly Weighted Neural Network?Vivek Ramanujan, Mitchell Wortsman, Aniruddha Kembhavi, Ali Farhadi et al.CVPR 2020
Related papers
- Why Random Pruning Is All We Need to Start SparseAdvait Harshal Gadhikar, Sohom Mukherjee, Rebekka BurkholzICML 2023 · 33 citations
- Masks, Signs, And Learning Rate RewindingAdvait Harshal Gadhikar, Rebekka BurkholzICLR 2024 · 15 citations
- Plant 'n' Seek: Can You Find the Winning Ticket?Jonas Fischer, Rebekka BurkholzICLR 2022 · 21 citations
- Sign-In to the Lottery: Reparameterizing Sparse TrainingAdvait Gadhikar, Tom Jacobs, Chao Zhou, Rebekka BurkholzNeurIPS 2025
- Lottery Ticket Preserves Weight Correlation: Is It Desirable or Not?Ning Liu, Geng Yuan, Zhengping Che, Xuan Shen et al.ICML 2021 · 34 citations
