Lune

AAAI2026Top-tier venue

AirWino: Optimized Winograd Convolution for Accelerating CNN Inference on ARMv8 Processors

Haoyuan Gui, Xiaoyu Zhang, Yifan Zhang, Ximeng Fu, Shiqi Sun, Leisheng Li, Huiyuan Li

2026Year

Abstract

As Convolutional Neural Networks (CNNs) continue to gain traction in deep learning, Winograd convolution has emerged as a key algorithm to enhance computational efficiency. Although ARM-based CPUs are increasingly prevalent in mobile devices, embedded systems and HPC servers, existing 2D Winograd convolution implementations for ARM often leave room for improvement in transformation efficiency, computational throughput, and overall versatility. Furthermore, the lack of tailored 3D Winograd convolution implementations for ARM architectures stems from the additional complexity of supporting higher-dimensional kernels. AirWino introduces a set of novel optimizations covering transformations, data layouts, micro-kernel computations, and parallelization strategies for both 2D and 3D Winograd convolution. It supports FP32 and FP16 precisions with filter sizes of 3 and 5, targeting a broad range of applications. Evaluations on four distinct ARM platforms show that AirWino consistently outperforms state-of-the-art libraries across various experimental scenarios and hardware configurations, highlighting its efficiency and portability.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 1574bfab-ca42-4f4f-85c4-aea634608ed3

Builds on3

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines
AirWino: Optimized Winograd Convolution for Accelerating CNN Inference on ARMv8 Processors | Lune Research