WinoTrain: Winograd-Aware Training for Accurate Full 8-bit Convolution Acceleration
Pierpaolo Morì, Shambhavi Balamuthu Sampath, Lukas Frickenstein, Manoj Rohit Vemparala, Nael Fasfous, Alexander Frickenstein, Walter Stechele, Claudio Passerone
摘要
Efficient inference is critical in realizing a low-power, real-time implementation of convolutional neural networks (CNNs) on compute and memory-constrained embedded platforms. Using quantization techniques and fast convolutional algorithms like Winograd, CNN inference can achieve benefits in latency and in energy consumption. Performing Winograd convolution involves (1) transforming the weights and activations to the Winograd domain, (2) performing element-wise multiplication on the transformed tensors, and (3) transforming the results back to the conventional spatial domain. Combining Winograd with quantization of all its steps results in severe accuracy degradation due to numerical instability. In this paper we propose a simple quantization-aware training technique, which quantizes all three steps of the Winograd convolution, while using a minimal number of scaling factors. Additionally, we propose an FPGA accelerator employing tiling and unrolling methods to highlight the performance benefits of using the full 8-bit quantized Winograd algorithm. We achieve 2× reduction in inference time compared to standard convolution on ResNet-18 for the ImageNet dataset, while improving the Top-1 accuracy by 55.7 p.p. compared to a standard post-training quantized Winograd variant of the network.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Towards Efficient and Accurate Winograd Convolution via Full QuantizationTianqi Chen, Weixiang Xu, Weihan Chen, Peisong Wang 等NeurIPS 2023 · 被引用 13 次
- Channel Balancing for Accurate Quantization of Winograd ConvolutionsVladimir Chikin, Vladimir KryzhanovskiyCVPR 2022 · 被引用 10 次
- Going Further With Winograd Convolutions: Tap-Wise Quantization for Efficient Inference on 4x4 TilesRenzo Andri, Beatrice Bussolino, Antonio Cipolletta, Lukas Cavigelli 等MICRO 2022 · 被引用 14 次
- SFC: Achieve Accurate Fast Convolution under Low-precision ArithmeticLiulu He, Yufei Zhao, Rui Gao, Yuan Du 等ICML 2024 · 被引用 3 次
- WINS: Winograd Structured Pruning for Fast Winograd ConvolutionCheonjun Park, Hyun Jae Oh, Mincheol Park, Hyunchan Moon 等ICCV 2025 · 被引用 2 次
