Accelerating winograd convolutions using symbolic computation and meta-programming
Arya Mazaheri, Tim Beringer, Matthew W. Moskewicz, Felix Wolf, Ali Jannesari
Abstract
Convolution operations are essential constituents of convolutional neural networks. Their efficient and performance-portable implementation demands tremendous programming effort and fine-tuning. Winograd's minimal filtering algorithm is a well-known method to reduce the computational complexity of convolution operations. Unfortunately, existing implementations of this algorithm are either vendor-specific or hard-coded to support a small subset of convolutions, thus limiting their versatility and performance portability. In this paper, we propose a novel method to optimize Winograd convolutions based on symbolic computation. Taking advantage meta-programming and auto-tuning, we further introduce a system to automate the generation of efficient and portable Winograd convolution code for various GPUs. We show that our optimization technique can effectively exploit repetitive patterns, enabling us to reduce the number of arithmetic operations by up to 62% without compromising numerical stability. Moreover, we demonstrate in experiments that we can generate efficient kernels with runtimes close to deep-learning libraries, requiring only a minimum of programming effort, which confirms the performance portability of our approach.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a069c033-5aa3-4fd3-b8c1-598532240703Related papers
- DWM: A Decomposable Winograd Method for Convolution AccelerationDi Huang, Xishan Zhang, Rui Zhang, Tian Zhi et al.AAAI 2020 · 31 citations
- WinoGen: A Highly Configurable Winograd Convolution IP Generator for Efficient CNN Acceleration on FPGAMingjun Li, Pengjia Li, Shuo Yin, Shixin Chen et al.DAC 2024 · 6 citations
- I/O lower bounds for auto-tuning of convolutions in CNNsXiaoyang Zhang, Junmin Xiao, Guangming TanPPoPP 2021 · 11 citations
- AirWino: Optimized Winograd Convolution for Accelerating CNN Inference on ARMv8 ProcessorsHaoyuan Gui, Xiaoyu Zhang, Yifan Zhang, Ximeng Fu et al.AAAI 2026
- DREW: Efficient Winograd CNN Inference with Deep ReuseRuofan Wu, Feng Zhang, Jiawei Guan, Zhen Zheng et al.WWW 2022 · 20 citations
