FlexHiNM-GP: Flexible Hierarchical Pruning via Region Allocation and Channel Permutation
Xiaodie Yi, Hayun Lee, Dongkun Shin
Abstract
N:M sparsity has emerged as a hardware-friendly pruning strategy, notably supported by NVIDIA’s Sparse Tensor Cores. While efficient, its fixed sparsity ratio restricts flexibility, making it difficult to adapt pruning granularity to varying weight importance across layers and architectures. To overcome this limitation, we propose FlexHiNM, a hybrid framework that adaptively partitions each layer into three regions: dense, vector-pruned, and N:M sparse, enabling finer-grained control while preserving hardware compatibility. To better preserve salient weights, we extend this to FlexHiNM-GP, which incorporates Gyro-Permutation, an iterative channel-rearrangement algorithm. Through successive sampling, clustering, and assignment, Gyro-Permutation aligns high-importance weights with structured sparsity patterns and mitigates suboptimal configurations in multi-level pruning. During gradual pruning, FlexHiNM-GP further employs a differentiable masking mechanism based on the Hard Concrete distribution, enabling gradient-based mask learning and preventing over-aggressive early pruning. Experiments on vision and language benchmarks demonstrate that FlexHiNM-GP consistently surpasses strong structured baselines and approaches the performance of unstructured pruning, validating the effectiveness of combining hybrid sparsity with learned masks and permutation strategies.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on13
- Movement Pruning: Adaptive Sparsity by Fine-TuningVictor Sanh, Thomas Wolf, Alexander M. RushNeurIPS 2020 · 656 citations
- Learning N: M Fine-grained Structured Sparse Neural Networks From ScratchAojun Zhou, Yukun Ma, Junnan Zhu, Jianbo Liu et al.ICLR 2021 · 301 citations
- Accelerated Sparse Neural Training: A Provable and Efficient Method to Find N: M Transposable MasksItay Hubara, Brian Chmiel, Moshe Island, Ron Banner et al.NeurIPS 2021 · 148 citations
- Channel Permutations for N: M SparsityJeff Pool, Chong YuNeurIPS 2021 · 75 citations
- MaskLLM: Learnable Semi-Structured Sparsity for Large Language ModelsGongfan Fang, Hongxu Yin, Saurav Muralidharan, Greg Heinrich et al.NeurIPS 2024 · 72 citations
Related papers
- BAME: Block-Aware Mask Evolution for Efficient N: M Sparse TrainingChenyi Yang, Wenjie Nie, Yuxin Zhang, Yuhang Wu et al.ICML 2025
- Learnable Permutation for Structured Sparsity on Transformer ModelsZekai Li, Ji Liu, Guanchen Li, Yixing Xu et al.AAAI 2026
- Bi-directional Masks for Efficient N: M Sparse TrainingYuxin Zhang, Yiting Luo, Mingbao Lin, Yunshan Zhong et al.ICML 2023 · 23 citations
- PermLLM: Learnable Channel Permutation for N: M Sparse Large Language ModelsLancheng Zou, Shuo Yin, Zehua Pei, Tsung-Yi Ho et al.NeurIPS 2025 · 1 citation
- MaxQ: Multi-Axis Query for N: m Sparsity NetworkJingyang Xiang, Siqi Li, Junhao Chen, Zhuangzhi Chen et al.CVPR 2024 · 2 citations
