ASBP: Automatic Structured Bit-Pruning for RRAM-based NN Accelerator
Songyun Qu, Bing Li, Ying Wang, Lei Zhang
摘要
Network sparsity or pruning is an extensively studied method to optimize the computation efficiency of deep neural networks (DNNs) for CMOS-based accelerators, such as FPGAs and GPUs. Though the RRAM-based accelerator has demonstrated superior performance and energy efficiency for DNN tasks, deploying the sparse neural networks desires dedicated consideration to save resource consumption without introducing the expensive index overhead and sophisticated control. To exploit the potential of sparse neural network design on the RRAM-based accelerator, we propose an automatic structured bit-pruning design, ASBP, to harmonize the optimization objective of DNN sparsity with efficient RRAM deployment. Specifically, ASBP prunes the bits of weight which are split into different crossbars and thus, free the zero-value crossbar when mapping the neural network into RRAM-based accelerators without extra hardware modification. Meanwhile, ASBP employs the reinforcement learning (RL) approach to automatically select the best crossbar-aware bit-sparsity strategy for any given neural network without laborious human efforts. According to our experiments on a set of representative neural networks, ASBP saves up to 79.01% energy consumption and 54.79% area overhead compared to the baseline that deploys the original DNN on the RRAM-based accelerator. Besides, ASBP outperforms the state-of-the-art bit-sparsity design by 1.4x in terms of the energy reduction on the RRAM-based accelerator.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- RAELLA: Reforming the Arithmetic for Efficient, Low-Resolution, and Low-Loss Analog PIM: No Retraining Required!Tanner Andrulis, Joel S. Emer, Vivienne SzeISCA 2023 · 被引用 45 次
- BBS: Bi-Directional Bit-Level Sparsity for Deep Learning AccelerationYuzong Chen, Jian Meng, Jae-sun Seo, Mohamed S. AbdelfattahMICRO 2024 · 被引用 25 次
相关 Paper
- BitPruner: Network Pruning for Bit-serial AcceleratorsXiandong Zhao, Ying Wang, Cheng Liu, Cong Shi 等DAC 2020 · 被引用 29 次
- RePIM: Joint Exploitation of Activation and Weight Repetitions for In-ReRAM DNN AccelerationChen-Yang Tsai, Chin-Fu Nien, Tz-Ching Yu, Hung-Yu Yeh 等DAC 2021 · 被引用 22 次
- Effective zero compression on ReRAM-based sparse DNN acceleratorsHoon Shin, Rihae Park, Seung Yul Lee, Yeonhong Park 等DAC 2022 · 被引用 11 次
- BitPattern: Enabling Efficient Bit-Serial Acceleration of Deep Neural Networks through Bit-Pattern PruningGang Wang, Siqi Cai, Zhenyu Li, Wenjie Li 等DAC 2025
- PattPIM: A Practical ReRAM-Based DNN Accelerator by Reusing Weight Pattern RepetitionsYuhao Zhang, Zhiping Jia, Yungang Pan, Hongchao Du 等DAC 2020 · 被引用 27 次
