I/O lower bounds for auto-tuning of convolutions in CNNs
Xiaoyang Zhang, Junmin Xiao, Guangming Tan
摘要
Convolution is the most time-consuming part in the computation of convolutional neural networks (CNNs), which have achieved great successes in numerous practical applications. Due to the complex data dependency and the increase in the amount of model samples, the convolution suffers from high overhead on data movement (i.e., memory access). This work provides comprehensive analysis and methodologies to minimize the communication for the convolution in CNNs. With an in-depth analysis of the recent I/O complexity theory under the red-blue game model, we develop a general I/O lower bound theory for a composite algorithm which consists of several different sub-computations. Based on the proposed theory, we establish the data movement lower bound results for two main convolution algorithms in CNNs, namely the direct convolution and Winograd algorithm, which represents the direct and indirect implementations of a convolution respectively. Next, derived from I/O lower bound results, we design the near I/O-optimal dataflow strategies for the two main convolution algorithms by fully exploiting the data reuse. Furthermore, in order to push the envelope of performance of the near I/O-optimal dataflow strategies further, an aggressive design of auto-tuning based on I/O lower bounds, is proposed to search an optimal parameter configuration for the direct convolution and Winograd algorithm on GPU, such as the number of threads and the size of shared memory used in each thread block. Finally, experiment evaluation results on the direct convolution and Winograd algorithm show that our dataflow strategies with the auto-tuning approach can achieve about 3.32× performance speedup on average over cuDNN. In addition, compared with TVM, which represents the state-of-the-art technique for auto-tuning, not only our auto-tuning method based on I/O lower bounds can find the optimal parameter configuration faster, but also our solution has higher performance than the optimal solution provided by TVM.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper3
- Optimizing batched Winograd convolution on GPUsDa Yan, Wei Wang, Xiaowen ChuPPoPP 2020 · 被引用 64 次
- Communication Lower Bound in Convolution AcceleratorsXiaoming Chen, Yinhe Han, Yu WangHPCA 2020 · 被引用 42 次
- Optimizing the Memory Hierarchy by Compositing Automatic Transformations on Computations and DataJie Zhao, Peng DiMICRO 2020 · 被引用 32 次
相关 Paper
- Accelerating winograd convolutions using symbolic computation and meta-programmingArya Mazaheri, Tim Beringer, Matthew W. Moskewicz, Felix Wolf 等EuroSys 2020 · 被引用 5 次
- Analytical characterization and design space exploration for optimization of CNNsRui Li, Yufan Xu, Aravind Sukumaran-Rajam, Atanas Rountev 等ASPLOS 2021 · 被引用 52 次
- DREW: Efficient Winograd CNN Inference with Deep ReuseRuofan Wu, Feng Zhang, Jiawei Guan, Zhen Zheng 等WWW 2022 · 被引用 20 次
- A History-Based Auto-Tuning Framework for Fast and High-Performance DNN Design on GPUJiandong Mu, Mengdi Wang, Lanbo Li, Jun Yang 等DAC 2020 · 被引用 15 次
- Optimizing Direct Convolutions on ARM Multi-CoresPengyu Wang, Weiling Yang, Jianbin Fang, Dezun Dong 等SC 2023 · 被引用 6 次
