Lune

DAC2023Top-tier venue

HBP: Hierarchically Balanced Pruning and Accelerator Co-Design for Efficient DNN Inference

Ao Ren, Yuhao Wang, Tao Zhang, Jiaxing Shi, Duo Liu, Xianzhang Chen, Yujuan Tan, Yuan Xie

2023Year
4Citations

Abstract

Weight pruning is studied to accelerate DNN inference by reducing the parameters and computations. Irregular pruning achieves high sparsity while incurring low computation parallelism and imbalanced workloads. The coarse-grained structured pruning sacrifices sparsity for higher parallelism. To strike a better balance, we propose Hierarchically Balanced Pruning by applying fine-grained but structured adjustments based on irregular pruning. Besides, it partitions the weight matrix into hierarchical blocks and constrains the sparsity of the blocks for balanced workloads. Furthermore, an accelerator is proposed to unleash the power of the pruning method. Experimental results show our method achieves 1.1×-6 higher sparsity than prior studies, and the accelerator achieves 1.2×-13× speedup and 3.3× energy efficiency improvement than its counterparts.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get 9dd8c7eb-6cd0-4ecb-a2d5-732df5adfa49

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines