Balanced Column-Wise Block Pruning for Maximizing GPU Parallelism
Cheonjun Park, Mincheol Park, Hyun Jae Oh, Minkyu Kim, Myung Kuk Yoon, Suhyun Kim, Won Woo Ro
摘要
Pruning has been an effective solution to reduce the number of computations and the memory requirement in deep learning. The pruning unit plays an important role in exploiting the GPU resources efficiently. The filter is proposed as a simple pruning unit of structured pruning. However, since the filter is quite large as pruning unit, the accuracy drop is considerable with a high pruning ratio. GPU rearranges the weight and input tensors into tiles (blocks) for efficient computation. To fully utilize GPU resources, this tile structure should be considered, which is the goal of block pruning. However, previous block pruning prunes both row vectors and column vectors. Pruning of row vectors in a tile corresponds to filter pruning, and it also interferes with column-wise block pruning of the following layer. In contrast, column vectors are much smaller than row vectors and can achieve lower accuracy drop. Additionally, if the pruning ratio for each tile is different, GPU utilization can be limited by imbalanced workloads by irregular-sized blocks. The same pruning ratio for the weight tiles processed in parallel enables the actual inference process to fully utilize the resources without idle time. This paper proposes balanced column-wise block pruning, named BCBP, to satisfy two conditions: the column-wise minimal size of the pruning unit and balanced workloads. We demonstrate that BCBP is superior to previous pruning methods through comprehensive experiments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- DEPrune: Depth-wise Separable Convolution Pruning for Maximizing GPU ParallelismCheonjun Park, Mincheol Park, Hyunchan Moon, Myung Kuk Yoon 等NeurIPS 2024 · 被引用 10 次
- REPrune: Channel Pruning via Kernel Representative SelectionMincheol Park, Dongjin Kim, Cheonjun Park, Yuna Park 等AAAI 2024 · 被引用 5 次
- WINS: Winograd Structured Pruning for Fast Winograd ConvolutionCheonjun Park, Hyun Jae Oh, Mincheol Park, Hyunchan Moon 等ICCV 2025 · 被引用 2 次
- FlexHiNM-GP: Flexible Hierarchical Pruning via Region Allocation and Channel PermutationXiaodie Yi, Hayun Lee, Dongkun ShinICLR 2026
它引用的顶会 Paper10
- Pruning neural networks without any data by iteratively conserving synaptic flowHidenori Tanaka, Daniel Kunin, Daniel L. K. Yamins, Surya GanguliNeurIPS 2020 · 被引用 884 次
- Learning N: M Fine-grained Structured Sparse Neural Networks From ScratchAojun Zhou, Yukun Ma, Junnan Zhu, Jianbo Liu 等ICLR 2021 · 被引用 301 次
- Sparse GPU kernels for deep learningTrevor Gale, Matei Zaharia, Cliff Young, Erich ElsenSC 2020 · 被引用 170 次
- Only Train Once: A One-Shot Neural Network Training And Pruning FrameworkTianyi Chen, Bo Ji, Tianyu Ding, Biyi Fang 等NeurIPS 2021 · 被引用 135 次
- YOLObile: Real-Time Object Detection on Mobile Devices via Compression-Compilation Co-DesignYuxuan Cai, Hongjia Li, Geng Yuan, Wei Niu 等AAAI 2021 · 被引用 124 次
相关 Paper
- Accelerating sparse DNN models without hardware-support via tile-wise sparsityCong Guo, Bo Yang Hsueh, Jingwen Leng, Yuxian Qiu 等SC 2020 · 被引用 65 次
- Pruning Filter in FilterFanxu Meng, Hao Cheng, Ke Li, Huixiang Luo 等NeurIPS 2020 · 被引用 130 次
- Shfl-BW: accelerating deep neural network inference with tensor-core aware weight pruningGuyue Huang, Haoran Li, Minghai Qin, Fei Sun 等DAC 2022 · 被引用 19 次
- HBP: Hierarchically Balanced Pruning and Accelerator Co-Design for Efficient DNN InferenceAo Ren, Yuhao Wang, Tao Zhang, Jiaxing Shi 等DAC 2023 · 被引用 4 次
- Taming Unstructured Sparsity on GPUs via Latency-Aware OptimizationMaohua Zhu, Yuan XieDAC 2020 · 被引用 3 次
