Communication Lower Bound in Convolution Accelerators
Xiaoming Chen, Yinhe Han, Yu Wang
Abstract
In current convolutional neural network (CNN) accelerators, communication (i.e., memory access) dominates the energy consumption. This work provides comprehensive analysis and methodologies to minimize the communication for CNN accelerators. For the off-chip communication, we derive the theoretical lower bound for any convolutional layer and propose a dataflow to reach the lower bound. This fundamental problem has never been solved by prior studies. The on-chip communication is minimized based on an elaborate workload and storage mapping scheme. We in addition design a communication-optimal CNN accelerator architecture. Evaluations based on the 65nm technology demonstrate that the proposed architecture nearly reaches the theoretical minimum communication in a three-level memory hierarchy and it is computation dominant. The gap between the energy efficiency of our accelerator and the theoretical best value is only 37-87%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- I/O lower bounds for auto-tuning of convolutions in CNNsXiaoyang Zhang, Junmin Xiao, Guangming TanPPoPP 2021 · 11 citations
- Cocco: Hardware-Mapping Co-Exploration towards Memory Capacity-Communication OptimizationZhanhong Tan, Zijian Zhu, Kaisheng MaASPLOS 2024 · 9 citations
Related papers
- Analytical characterization and design space exploration for optimization of CNNsRui Li, Yufan Xu, Aravind Sukumaran-Rajam, Atanas Rountev et al.ASPLOS 2021 · 52 citations
- DRMap: A Generic DRAM Data Mapping Policy for Energy-Efficient Processing of Convolutional Neural NetworksRachmad Vidya Wicaksana Putra, Muhammad Abdullah Hanif, Muhammad ShafiqueDAC 2020 · 36 citations
- GoSPA: An Energy-efficient High-performance Globally Optimized SParse Convolutional Neural Network AcceleratorChunhua Deng, Yang Sui, Siyu Liao, Xuehai Qian et al.ISCA 2021 · 77 citations
- QuiltNet: efficient deep learning inference on multi-chip accelerators using model partitioningJongho Park, Hyukjun Kwon, Seowoo Kim, Junyoung Lee et al.DAC 2022 · 6 citations
- STC: Significance-aware Transform-based Codec Framework for External Memory Access ReductionFeng Xiong, Fengbin Tu, Man Shi, Yang Wang et al.DAC 2020 · 23 citations
