A Convolution Neural Network Accelerator Design with Weight Mapping and Pipeline Optimization
Lixia Han, Peng Huang, Zheng Zhou, Yiyang Chen, Xiaoyan Liu, Jinfeng Kang
Abstract
The pipeline is an efficient solution to boost performance in non-volatile memory based computing in memory (nvCIM) convolution neural network (CNN) accelerators. However, the previous works seldom focus on pipeline optimization from the perspective of the whole system, especially overlooking the effect of buffer access. In this work, we propose a high-performance NVM-based CNN accelerator with a balanced pipeline design, which takes account of both the macro computing and the buffer access. At the operator level, a matrix-based weight mapping method is proposed to reduce buffer access delay. At the macro level, decoupled access and execution design is introduced to shorten the single-layer latency. At the system level, a hybrid inter/intra-tile design is presented to balance the overall latency across CNN layers. With the collaboration among three methods, we construct a well-balanced pipeline for the nvCIM accelerator at a smaller hardware cost. Experiments show that our pipeline design can achieve 3.7х, 7.5х, and 3.5х throughput improvement for recognition of ImageNet with ResNet18, VGG19, and ResNet34 models, respectively.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- PIMCOMP: A Universal Compilation Framework for Crossbar-based PIM DNN AcceleratorsXiaotian Sun, Xinyu Wang, Wanqian Li, Lei Wang et al.DAC 2023 · 19 citations
- TransPIM: A Memory-based Acceleration via Software-Hardware Co-Design for TransformerMinxuan Zhou, Weihong Xu, Jaeyoung Kang, Tajana RosingHPCA 2022 · 142 citations
- DRMap: A Generic DRAM Data Mapping Policy for Energy-Efficient Processing of Convolutional Neural NetworksRachmad Vidya Wicaksana Putra, Muhammad Abdullah Hanif, Muhammad ShafiqueDAC 2020 · 36 citations
- EPIM: Efficient Processing-In-Memory Accelerators based on EpitomeChenyu Wang, Zhen Dong, Daquan Zhou, Zhenhua Zhu et al.DAC 2024
- Accelerating Sparse Attention with a Reconfigurable Non-volatile Processing-In-Memory ArchitectureQilin Zheng, Shiyu Li, Yitu Wang, Ziru Li et al.DAC 2023 · 14 citations
