Going Further With Winograd Convolutions: Tap-Wise Quantization for Efficient Inference on 4x4 Tiles
Renzo Andri, Beatrice Bussolino, Antonio Cipolletta, Lukas Cavigelli, Zhe Wang
Abstract
Most of today's computer vision pipelines are built around deep neural networks, where convolution operations require most of the generally high compute effort. The Winograd convolution algorithm computes convolutions with fewer multiply-accumulate operations (MACs) compared to the standard algorithm, reducing the operation count by a factor of 2.25× for 3×3 convolutions when using the version with 2×2sized tiles F 2 . Even though the gain is significant, the Winograd algorithm with larger tile sizes, i.e., F 4 , offers even more potential in improving throughput and energy efficiency, as it reduces the required MACs by 4×. Unfortunately, the Winograd algorithm with larger tile sizes introduces numerical issues that prevent its use on integer domain-specific accelerators (DSAs) and higher computational overhead to transform input and output data between spatial and Winograd domains.
To unlock the full potential of Winograd F 4 , we propose a novel tap-wise quantization method that overcomes the numerical issues of using larger tiles, enabling integer-only inference. Moreover, we present custom hardware units that process the Winograd transformations in a power-and area-efficient way, and we show how to integrate such custom modules in an industrial-grade, programmable DSA. An extensive experimental evaluation on a large set of state-of-the-art computer vision benchmarks reveals that the tap-wise quantization algorithm makes the quantized Winograd F 4 network almost as accurate as the FP32 baseline. The Winograd-enhanced DSA achieves up to 1.85× gain in energy efficiency and up to 1.83× end-toend speed-up for state-of-the-art segmentation and detection networks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 35acc3c6-798c-4f2e-83ee-1622e601f508Cited by top-tier papers3
- Towards Efficient and Accurate Winograd Convolution via Full QuantizationTianqi Chen, Weixiang Xu, Weihan Chen, Peisong Wang et al.NeurIPS 2023 · 13 citations
- SFC: Achieve Accurate Fast Convolution under Low-precision ArithmeticLiulu He, Yufei Zhao, Rui Gao, Yuan Du et al.ICML 2024 · 3 citations
- WINS: Winograd Structured Pruning for Fast Winograd ConvolutionCheonjun Park, Hyun Jae Oh, Mincheol Park, Hyunchan Moon et al.ICCV 2025 · 2 citations
Builds on2
Related papers
- WinoTrain: Winograd-Aware Training for Accurate Full 8-bit Convolution AccelerationPierpaolo Morì, Shambhavi Balamuthu Sampath, Lukas Frickenstein, Manoj Rohit Vemparala et al.DAC 2023 · 5 citations
- WinoGen: A Highly Configurable Winograd Convolution IP Generator for Efficient CNN Acceleration on FPGAMingjun Li, Pengjia Li, Shuo Yin, Shixin Chen et al.DAC 2024 · 6 citations
- Channel Balancing for Accurate Quantization of Winograd ConvolutionsVladimir Chikin, Vladimir KryzhanovskiyCVPR 2022 · 10 citations
- Optimizing batched Winograd convolution on GPUsDa Yan, Wei Wang, Xiaowen ChuPPoPP 2020 · 64 citations
- DWM: A Decomposable Winograd Method for Convolution AccelerationDi Huang, Xishan Zhang, Rui Zhang, Tian Zhi et al.AAAI 2020 · 31 citations
