NTK-SAP: Improving neural network pruning by aligning training dynamics
Yite Wang, Dawei Li, Ruoyu Sun
Abstract
Pruning neural networks before training has received increasing interest due to its potential to reduce training time and memory. One popular method is to prune the connections based on a certain metric, but it is not entirely clear what metric is the best choice. Recent advances in neural tangent kernel (NTK) theory suggest that the training dynamics of large enough neural networks is closely related to the spectrum of the NTK. Motivated by this finding, we propose to prune the connections that have the least influence on the spectrum of the NTK. This method can help maintain the NTK spectrum, which may help align the training dynamics to that of its dense counterpart. However, one possible issue is that the fixedweight-NTK corresponding to a given initial point can be very different from the NTK corresponding to later iterates during the training phase. We further propose to sample multiple realizations of random weights to estimate the NTK spectrum. Note that our approach is weight-agnostic, which is different from most existing methods that are weight-dependent. In addition, we use random inputs to compute the fixed-weight-NTK, making our method data-agnostic as well. We name our foresight pruning algorithm Neural Tangent Kernel Spectrum-Aware Pruning (NTK-SAP). Empirically, our method achieves better performance than all baselines on multiple datasets. Our code is available at https://github. com/YiteWang/NTK-SAP .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext da811a9e-dc67-4038-8aad-5f45bff2a272Cited by top-tier papers14
- SwitchTab: Switched Autoencoders Are Effective Tabular LearnersJing Wu, Suiyao Chen, Qi Zhao, Renat Sergazinov et al.AAAI 2024 · 60 citations
- LEMON: Lossless model expansionYite Wang, Jiahao Su, Hanlin Lu, Cong Xie et al.ICLR 2024 · 25 citations
- Balanced Training for Sparse GANsYite Wang, Jing Wu, Naira Hovakimyan, Ruoyu SunNeurIPS 2023 · 16 citations
- Train Faster, Perform Better: Modular Adaptive Training in Over-Parameterized ModelsYubin Shi, Yixuan Chen, Mingzhi Dong, Xiaochen Yang et al.NeurIPS 2023 · 5 citations
- SepPrune: Structured Pruning for Efficient Deep Speech SeparationYuqi Li, Kai Li, Xin Yin, Zhifei Yang et al.AAAI 2026 · 4 citations
Builds on27
- Pruning neural networks without any data by iteratively conserving synaptic flowHidenori Tanaka, Daniel Kunin, Daniel L. K. Yamins, Surya GanguliNeurIPS 2020 · 884 citations
- Linear Mode Connectivity and the Lottery Ticket HypothesisJonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, Michael CarbinICML 2020 · 750 citations
- Picking Winning Tickets Before Training by Preserving Gradient FlowChaoqi Wang, Guodong Zhang, Roger B. GrosseICLR 2020 · 743 citations
- Rigging the Lottery: Making All Tickets WinnersUtku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro et al.ICML 2020 · 723 citations
- Comparing Rewinding and Fine-tuning in Neural Network PruningAlex Renda, Jonathan Frankle, Michael CarbinICLR 2020 · 437 citations
Related papers
- Finding Lottery Tickets in Vision Models via Data-Driven Spectral Foresight PruningLeonardo Iurada, Marco Ciccone, Tatiana TommasiCVPR 2024
- Neural Architecture Search on ImageNet in Four GPU Hours: A Theoretically Inspired PerspectiveWuyang Chen, Xinyu Gong, Zhangyang WangICLR 2021 · 51 citations
- The Surprising Effectiveness of Infinite-Width NTKs for Characterizing and Improving Model TrainingJoshua DeOliveira, Walter Gerych, Elke A. RundensteinerAAAI 2025 · 1 citation
- The Graphon Limit Hypothesis: Understanding Neural Network Pruning via Infinite Width AnalysisHoang Pham, The Anh Ta, Tom Jacobs, Rebekka Burkholz et al.NeurIPS 2025 · 2 citations
- PHEW : Constructing Sparse Networks that Learn Fast and Generalize Well without Training DataShreyas Malakarjun Patil, Constantine DovrolisICML 2021 · 26 citations
