Sign-In to the Lottery: Reparameterizing Sparse Training
Advait Gadhikar, Tom Jacobs, Chao Zhou, Rebekka Burkholz
摘要
The performance gap between training sparse neural networks from scratch (PaI) and dense-to-sparse training, presents a major roadblock for efficient deep learning. According to the Lottery Ticket Hypothesis, PaI hinges on finding a problem specific parameter initialization, given a sparse mask. As we show, to this end, determining correct parameter signs is sufficient. Yet, they remain elusive to PaI. To address this issue, we propose Sign-In, which employs a dynamic reparameterization that provably induces sign flips. Such sign flips are complementary to the ones that dense-to-sparse training can accomplish, rendering Sign-In as an orthogonal method. While our experiments and theory suggest performance improvements of PaI, they also carve out the main open challenge to close the gap between PaI and dense-to-sparse training.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Hyperbolic Aware Minimization: Implicit Bias for SparsityTom Jacobs, Advait Gadhikar, Celia Rubio-Madrigal, Rebekka BurkholzICLR 2026 · 被引用 3 次
- SparseOpt: Addressing Normalization-induced Gradient Skew in Sparse TrainingAdnan Mohammed, Rohan Jain, Tom Jacobs, Ekansh Sharma 等ICML 2026
- HASTE: Hardware-Aware Dynamic Sparse Training for Large Output SpacesNasib Ullah, Jinbin Zhang, Jean Lucien Randrianantenaina, Erik Schultheis 等ICML 2026
- Theoretical Analysis of Sparse Optimization with Reparameterization, Weight Decay, and Adaptive Learning RateHuangyu Xu, Jingqin Yang, Qianqian Xu, Jiaye TengICML 2026
它引用的顶会 Paper40
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Pruning neural networks without any data by iteratively conserving synaptic flowHidenori Tanaka, Daniel Kunin, Daniel L. K. Yamins, Surya GanguliNeurIPS 2020 · 被引用 884 次
- Picking Winning Tickets Before Training by Preserving Gradient FlowChaoqi Wang, Guodong Zhang, Roger B. GrosseICLR 2020 · 被引用 743 次
- Rigging the Lottery: Making All Tickets WinnersUtku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro 等ICML 2020 · 被引用 723 次
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan 等ICLR 2020 · 被引用 705 次
相关 Paper
- Gradient Flow in Sparse Neural Networks and How Lottery Tickets WinUtku Evci, Yani Ioannou, Cem Keskin, Yann N. DauphinAAAI 2022 · 被引用 106 次
- Sparse Training from Random Initialization: Aligning Lottery Ticket Masks using Weight SymmetryMohammed Adnan, Rohan Jain, Ekansh Sharma, Rahul Krishnan 等ICML 2025
- Masks, Signs, And Learning Rate RewindingAdvait Harshal Gadhikar, Rebekka BurkholzICLR 2024 · 被引用 15 次
- Find A Winning Sign: Sign Is All We Need to Win the LotteryJunghun Oh, Sungyong Baik, Kyoung Mu LeeICLR 2025
- Winning the Lottery with Continuous SparsificationPedro Savarese, Hugo Silva, Michael MaireNeurIPS 2020 · 被引用 162 次
