FlexHiNM-GP: Flexible Hierarchical Pruning via Region Allocation and Channel Permutation
Xiaodie Yi, Hayun Lee, Dongkun Shin
摘要
N:M sparsity has emerged as a hardware-friendly pruning strategy, notably supported by NVIDIA’s Sparse Tensor Cores. While efficient, its fixed sparsity ratio restricts flexibility, making it difficult to adapt pruning granularity to varying weight importance across layers and architectures. To overcome this limitation, we propose FlexHiNM, a hybrid framework that adaptively partitions each layer into three regions: dense, vector-pruned, and N:M sparse, enabling finer-grained control while preserving hardware compatibility. To better preserve salient weights, we extend this to FlexHiNM-GP, which incorporates Gyro-Permutation, an iterative channel-rearrangement algorithm. Through successive sampling, clustering, and assignment, Gyro-Permutation aligns high-importance weights with structured sparsity patterns and mitigates suboptimal configurations in multi-level pruning. During gradual pruning, FlexHiNM-GP further employs a differentiable masking mechanism based on the Hard Concrete distribution, enabling gradient-based mask learning and preventing over-aggressive early pruning. Experiments on vision and language benchmarks demonstrate that FlexHiNM-GP consistently surpasses strong structured baselines and approaches the performance of unstructured pruning, validating the effectiveness of combining hybrid sparsity with learned masks and permutation strategies.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- Movement Pruning: Adaptive Sparsity by Fine-TuningVictor Sanh, Thomas Wolf, Alexander M. RushNeurIPS 2020 · 被引用 656 次
- Learning N: M Fine-grained Structured Sparse Neural Networks From ScratchAojun Zhou, Yukun Ma, Junnan Zhu, Jianbo Liu 等ICLR 2021 · 被引用 301 次
- Accelerated Sparse Neural Training: A Provable and Efficient Method to Find N: M Transposable MasksItay Hubara, Brian Chmiel, Moshe Island, Ron Banner 等NeurIPS 2021 · 被引用 148 次
- Channel Permutations for N: M SparsityJeff Pool, Chong YuNeurIPS 2021 · 被引用 75 次
- MaskLLM: Learnable Semi-Structured Sparsity for Large Language ModelsGongfan Fang, Hongxu Yin, Saurav Muralidharan, Greg Heinrich 等NeurIPS 2024 · 被引用 72 次
相关 Paper
- BAME: Block-Aware Mask Evolution for Efficient N: M Sparse TrainingChenyi Yang, Wenjie Nie, Yuxin Zhang, Yuhang Wu 等ICML 2025
- Learnable Permutation for Structured Sparsity on Transformer ModelsZekai Li, Ji Liu, Guanchen Li, Yixing Xu 等AAAI 2026
- Bi-directional Masks for Efficient N: M Sparse TrainingYuxin Zhang, Yiting Luo, Mingbao Lin, Yunshan Zhong 等ICML 2023 · 被引用 23 次
- PermLLM: Learnable Channel Permutation for N: M Sparse Large Language ModelsLancheng Zou, Shuo Yin, Zehua Pei, Tsung-Yi Ho 等NeurIPS 2025 · 被引用 1 次
- MaxQ: Multi-Axis Query for N: m Sparsity NetworkJingyang Xiang, Siqi Li, Junhao Chen, Zhuangzhi Chen 等CVPR 2024 · 被引用 2 次
