Going Further With Winograd Convolutions: Tap-Wise Quantization for Efficient Inference on 4x4 Tiles
Renzo Andri, Beatrice Bussolino, Antonio Cipolletta, Lukas Cavigelli, Zhe Wang
摘要
Most of today's computer vision pipelines are built around deep neural networks, where convolution operations require most of the generally high compute effort. The Winograd convolution algorithm computes convolutions with fewer multiply-accumulate operations (MACs) compared to the standard algorithm, reducing the operation count by a factor of 2.25× for 3×3 convolutions when using the version with 2×2sized tiles F 2 . Even though the gain is significant, the Winograd algorithm with larger tile sizes, i.e., F 4 , offers even more potential in improving throughput and energy efficiency, as it reduces the required MACs by 4×. Unfortunately, the Winograd algorithm with larger tile sizes introduces numerical issues that prevent its use on integer domain-specific accelerators (DSAs) and higher computational overhead to transform input and output data between spatial and Winograd domains.
To unlock the full potential of Winograd F 4 , we propose a novel tap-wise quantization method that overcomes the numerical issues of using larger tiles, enabling integer-only inference. Moreover, we present custom hardware units that process the Winograd transformations in a power-and area-efficient way, and we show how to integrate such custom modules in an industrial-grade, programmable DSA. An extensive experimental evaluation on a large set of state-of-the-art computer vision benchmarks reveals that the tap-wise quantization algorithm makes the quantized Winograd F 4 network almost as accurate as the FP32 baseline. The Winograd-enhanced DSA achieves up to 1.85× gain in energy efficiency and up to 1.83× end-toend speed-up for state-of-the-art segmentation and detection networks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Towards Efficient and Accurate Winograd Convolution via Full QuantizationTianqi Chen, Weixiang Xu, Weihan Chen, Peisong Wang 等NeurIPS 2023 · 被引用 13 次
- SFC: Achieve Accurate Fast Convolution under Low-precision ArithmeticLiulu He, Yufei Zhao, Rui Gao, Yuan Du 等ICML 2024 · 被引用 3 次
- WINS: Winograd Structured Pruning for Fast Winograd ConvolutionCheonjun Park, Hyun Jae Oh, Mincheol Park, Hyunchan Moon 等ICCV 2025 · 被引用 2 次
它引用的顶会 Paper2
相关 Paper
- WinoTrain: Winograd-Aware Training for Accurate Full 8-bit Convolution AccelerationPierpaolo Morì, Shambhavi Balamuthu Sampath, Lukas Frickenstein, Manoj Rohit Vemparala 等DAC 2023 · 被引用 5 次
- WinoGen: A Highly Configurable Winograd Convolution IP Generator for Efficient CNN Acceleration on FPGAMingjun Li, Pengjia Li, Shuo Yin, Shixin Chen 等DAC 2024 · 被引用 6 次
- Channel Balancing for Accurate Quantization of Winograd ConvolutionsVladimir Chikin, Vladimir KryzhanovskiyCVPR 2022 · 被引用 10 次
- Optimizing batched Winograd convolution on GPUsDa Yan, Wei Wang, Xiaowen ChuPPoPP 2020 · 被引用 64 次
- DWM: A Decomposable Winograd Method for Convolution AccelerationDi Huang, Xishan Zhang, Rui Zhang, Tian Zhi 等AAAI 2020 · 被引用 31 次
