Optimizing batched Winograd convolution on GPUs
Da Yan, Wei Wang, Xiaowen Chu
摘要
In this paper, we present an optimized implementation for single-precision Winograd convolution on NVIDIA Volta and Turing GPUs. Compared with the state-of-the-art Winograd convolution in cuDNN 7.6.1, our implementation achieves up to 2.13× speedup on Volta V100 and up to 2.65× speedup on Turing RTX2070. On both Volta and Turing GPUs, our implementation achieves up to 93% of device peak.
Apart from analyzing and benchmarking different highlevel optimization options, we also build a SASS assembler TuringAs for Volta and Turing that enables tuning the performance at the native assembly level. The new optimization opportunities uncovered by TuringAs not only improve the Winograd convolution but can also benefit CUDA compilers and native assembly programming. We have released TuringAs as an open-source software. To the best of our knowledge, this is the first public-available assembler for Volta and Turing GPUs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- DTC-SpMM: Bridging the Gap in Accelerating General Sparse Matrix Multiplication with Tensor CoresRuibo Fan, Wei Wang, Xiaowen ChuASPLOS 2024 · 被引用 46 次
- Causes and Effects of Unanticipated Numerical Deviations in Neural Network Inference FrameworksAlexander Schlögl, Nora Hofer, Rainer BöhmeNeurIPS 2023 · 被引用 31 次
- Morpheus: Extending the Last Level Cache Capacity in GPU Systems Using Idle GPU Core ResourcesSina Darabi, Mohammad Sadrosadati, Negar Akbarzadeh, Joël Lindegger 等MICRO 2022 · 被引用 24 次
- DREW: Efficient Winograd CNN Inference with Deep ReuseRuofan Wu, Feng Zhang, Jiawei Guan, Zhen Zheng 等WWW 2022 · 被引用 20 次
- I/O lower bounds for auto-tuning of convolutions in CNNsXiaoyang Zhang, Junmin Xiao, Guangming TanPPoPP 2021 · 被引用 11 次
相关 Paper
- GPNPU: Enabling Efficient Hardware-Based Direct Convolution with Multi-Precision Support in GPU Tensor CoresZhuoran Song, Jianfei Wang, Tianjian Li, Li Jiang 等DAC 2020 · 被引用 12 次
- Accelerating winograd convolutions using symbolic computation and meta-programmingArya Mazaheri, Tim Beringer, Matthew W. Moskewicz, Felix Wolf 等EuroSys 2020 · 被引用 5 次
- Going Further With Winograd Convolutions: Tap-Wise Quantization for Efficient Inference on 4x4 TilesRenzo Andri, Beatrice Bussolino, Antonio Cipolletta, Lukas Cavigelli 等MICRO 2022 · 被引用 14 次
- AirWino: Optimized Winograd Convolution for Accelerating CNN Inference on ARMv8 ProcessorsHaoyuan Gui, Xiaoyu Zhang, Yifan Zhang, Ximeng Fu 等AAAI 2026
- WINS: Winograd Structured Pruning for Fast Winograd ConvolutionCheonjun Park, Hyun Jae Oh, Mincheol Park, Hyunchan Moon 等ICCV 2025 · 被引用 2 次
