PIM-Prune: Fine-Grain DCNN Pruning for Crossbar-Based Process-In-Memory Architecture
Chaoqun Chu, Yanzhi Wang, Yilong Zhao, Xiaolong Ma, Shaokai Ye, Yunyan Hong, Xiaoyao Liang, Yinhe Han, Li Jiang
摘要
Deep Convolution Neural network (DCNN) pruning is an efficient way to reduce the resource and power consumption in a DCNN accelerator. Exploiting the sparsity in the weight matrices of DCNNs, however, is nontrivial if we deploy these DC-NNs in a crossbar-based Process-In-Memory (PIM) architecture, because of the crossbar structure. Structural pruning-exploiting a coarse-grained sparsity, such as filter/channel-level pruning-can result in a compressed weight matrix that fits the crossbar structure. However, this pruning method inevitably degrades the model accuracy. To solve this problem, in this paper, we propose PIM-PRUNE to exploit the finer-grained sparsity in PIM-architecture, and the resulting compressed weight matrices can significantly reduce the demand of crossbars with negligible accuracy loss. Further, we explore the design space of the crossbar, such as the crossbar size and aspect-ratio, from a new point-of-view of resource-oriented pruning. We find a trade-off existing between the pruning algorithm and the hardware overhead: a PIM with smaller crossbars is more friendly for pruning methods; however, the resulting peripheral circuit cause higher power consumption. Given a specific DCNN, we can suggest a sweet-spot of crossbar design to the optimal overall energy efficiency. Experimental results show that the proposed pruning method applied on Resnet18 can achieve up to 24.85× and 3.56× higher compression rate of occupied crossbars on CifarlO and Imagenet, respectively; while the accuracy loss is negligible, which is 4.56× and 1.99× better than the state-of-art methods.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper3
- RAELLA: Reforming the Arithmetic for Efficient, Low-Resolution, and Low-Loss Analog PIM: No Retraining Required!Tanner Andrulis, Joel S. Emer, Vivienne SzeISCA 2023 · 被引用 45 次
- MIME: adapting a single neural network for multi-task inference with memory-efficient dynamic pruningAbhiroop Bhattacharjee, Yeshwanth Venkatesha, Abhishek Moitra, Priyadarshini PandaDAC 2022 · 被引用 7 次
- EPIM: Efficient Processing-In-Memory Accelerators based on EpitomeChenyu Wang, Zhen Dong, Daquan Zhou, Zhenhua Zhu 等DAC 2024
相关 Paper
- DEPrune: Depth-wise Separable Convolution Pruning for Maximizing GPU ParallelismCheonjun Park, Mincheol Park, Hyunchan Moon, Myung Kuk Yoon 等NeurIPS 2024 · 被引用 10 次
- HBP: Hierarchically Balanced Pruning and Accelerator Co-Design for Efficient DNN InferenceAo Ren, Yuhao Wang, Tao Zhang, Jiaxing Shi 等DAC 2023 · 被引用 4 次
- PattPIM: A Practical ReRAM-Based DNN Accelerator by Reusing Weight Pattern RepetitionsYuhao Zhang, Zhiping Jia, Yungang Pan, Hongchao Du 等DAC 2020 · 被引用 27 次
- PCNN: Pattern-based Fine-Grained Regular Pruning Towards Optimizing CNN AcceleratorsZhanhong Tan, Jiebo Song, Xiaolong Ma, Sia Huat Tan 等DAC 2020 · 被引用 28 次
- Non-uniform DNN Structured Subnets Sampling for Dynamic InferenceLi Yang, Zhezhi He, Yu Cao, Deliang FanDAC 2020 · 被引用 12 次
