Channel Permutations for N: M Sparsity
Jeff Pool, Chong Yu
Abstract
We introduce channel permutations as a method to maximize the accuracy of N:M sparse networks. N:M sparsity requires N out of M consecutive elements to be zero and has been shown to maintain accuracy for many models and tasks with a simple prune and fine-tune workflow. By permuting weight matrices along their channel dimension and adjusting the surrounding layers appropriately, we demonstrate accuracy recovery for even small, parameter-efficient networks, without affecting inference run-time. We also present both a quality metric to simplify judging permutations as well as efficient methods to search for high-quality permutations, including two optimizations to escape local minima. Finally, we share an ablation study to show the importance of each part of our search algorithm, experimental results showing correlation between our quality metric and final network accuracy, improved sparse network accuracy using our techniques with insignificant overhead to training time, and the transformation of unstructured to structured sparse workloads. Code to use these techniques when generating a 2:4
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7b1479fd-73e2-44da-ab1a-0e3289da97d2Cited by top-tier papers26
- Plug-and-Play: An Efficient Post-training Pruning Method for Large Language ModelsYingtao Zhang, Haoli Bai, Haokun Lin, Jialin Zhao et al.ICLR 2024 · 72 citations
- MaskLLM: Learnable Semi-Structured Sparsity for Large Language ModelsGongfan Fang, Hongxu Yin, Saurav Muralidharan, Greg Heinrich et al.NeurIPS 2024 · 72 citations
- Learning Best Combination for Efficient N: M SparsityYuxin Zhang, Mingbao Lin, Zhihang Lin, Yiting Luo et al.NeurIPS 2022 · 66 citations
- Scaling Laws for Sparsely-Connected Foundation ModelsElias Frantar, Carlos Riquelme Ruiz, Neil Houlsby, Dan Alistarh et al.ICLR 2024 · 48 citations
- SPDY: Accurate Pruning with Speedup GuaranteesElias Frantar, Dan AlistarhICML 2022 · 45 citations
Builds on7
- MetaPruning: Meta Learning for Automatic Neural Network Channel PruningZechun Liu, Haoyuan Mu, Xiangyu Zhang, Zichao Guo et al.ICCV 2019 · 633 citations
- Comparing Rewinding and Fine-tuning in Neural Network PruningAlex Renda, Jonathan Frankle, Michael CarbinICLR 2020 · 437 citations
- Learning N: M Fine-grained Structured Sparse Neural Networks From ScratchAojun Zhou, Yukun Ma, Junnan Zhu, Jianbo Liu et al.ICLR 2021 · 301 citations
- Accelerated Sparse Neural Training: A Provable and Efficient Method to Find N: M Transposable MasksItay Hubara, Brian Chmiel, Moshe Island, Ron Banner et al.NeurIPS 2021 · 148 citations
- Tight Compression: Compressing CNN Model Tightly Through Unstructured Pruning and Simulated Annealing Based PermutationXizi Chen, Jingyang Zhu, Jingbo Jiang, Chi-Ying TsuiDAC 2020 · 30 citations
Related papers
- PermLLM: Learnable Channel Permutation for N: M Sparse Large Language ModelsLancheng Zou, Shuo Yin, Zehua Pei, Tsung-Yi Ho et al.NeurIPS 2025 · 1 citation
- DominoSearch: Find layer-wise fine-grained N: M sparse schemes from dense neural networksWei Sun, Aojun Zhou, Sander Stuijk, Rob G. J. Wijnhoven et al.NeurIPS 2021 · 67 citations
- BAME: Block-Aware Mask Evolution for Efficient N: M Sparse TrainingChenyi Yang, Wenjie Nie, Yuxin Zhang, Yuhang Wu et al.ICML 2025
- MaxQ: Multi-Axis Query for N: m Sparsity NetworkJingyang Xiang, Siqi Li, Junhao Chen, Zhuangzhi Chen et al.CVPR 2024 · 2 citations
- Bi-directional Masks for Efficient N: M Sparse TrainingYuxin Zhang, Yiting Luo, Mingbao Lin, Yunshan Zhong et al.ICML 2023 · 23 citations
