Optimizing batched Winograd convolution on GPUs
Da Yan, Wei Wang, Xiaowen Chu
Abstract
In this paper, we present an optimized implementation for single-precision Winograd convolution on NVIDIA Volta and Turing GPUs. Compared with the state-of-the-art Winograd convolution in cuDNN 7.6.1, our implementation achieves up to 2.13× speedup on Volta V100 and up to 2.65× speedup on Turing RTX2070. On both Volta and Turing GPUs, our implementation achieves up to 93% of device peak.
Apart from analyzing and benchmarking different highlevel optimization options, we also build a SASS assembler TuringAs for Volta and Turing that enables tuning the performance at the native assembly level. The new optimization opportunities uncovered by TuringAs not only improve the Winograd convolution but can also benefit CUDA compilers and native assembly programming. We have released TuringAs as an open-source software. To the best of our knowledge, this is the first public-available assembler for Volta and Turing GPUs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5a1295c5-891e-4f91-8a66-d15b62da82fdCited by top-tier papers9
- DTC-SpMM: Bridging the Gap in Accelerating General Sparse Matrix Multiplication with Tensor CoresRuibo Fan, Wei Wang, Xiaowen ChuASPLOS 2024 · 46 citations
- Causes and Effects of Unanticipated Numerical Deviations in Neural Network Inference FrameworksAlexander Schlögl, Nora Hofer, Rainer BöhmeNeurIPS 2023 · 31 citations
- Morpheus: Extending the Last Level Cache Capacity in GPU Systems Using Idle GPU Core ResourcesSina Darabi, Mohammad Sadrosadati, Negar Akbarzadeh, Joël Lindegger et al.MICRO 2022 · 24 citations
- DREW: Efficient Winograd CNN Inference with Deep ReuseRuofan Wu, Feng Zhang, Jiawei Guan, Zhen Zheng et al.WWW 2022 · 20 citations
- I/O lower bounds for auto-tuning of convolutions in CNNsXiaoyang Zhang, Junmin Xiao, Guangming TanPPoPP 2021 · 11 citations
Related papers
- GPNPU: Enabling Efficient Hardware-Based Direct Convolution with Multi-Precision Support in GPU Tensor CoresZhuoran Song, Jianfei Wang, Tianjian Li, Li Jiang et al.DAC 2020 · 12 citations
- Accelerating winograd convolutions using symbolic computation and meta-programmingArya Mazaheri, Tim Beringer, Matthew W. Moskewicz, Felix Wolf et al.EuroSys 2020 · 5 citations
- Going Further With Winograd Convolutions: Tap-Wise Quantization for Efficient Inference on 4x4 TilesRenzo Andri, Beatrice Bussolino, Antonio Cipolletta, Lukas Cavigelli et al.MICRO 2022 · 14 citations
- AirWino: Optimized Winograd Convolution for Accelerating CNN Inference on ARMv8 ProcessorsHaoyuan Gui, Xiaoyu Zhang, Yifan Zhang, Ximeng Fu et al.AAAI 2026
- WINS: Winograd Structured Pruning for Fast Winograd ConvolutionCheonjun Park, Hyun Jae Oh, Mincheol Park, Hyunchan Moon et al.ICCV 2025 · 2 citations
