WINS: Winograd Structured Pruning for Fast Winograd Convolution
Cheonjun Park, Hyun Jae Oh, Mincheol Park, Hyunchan Moon, Minsik Kim, Suhyun Kim, Myung Kuk Yoon, Won Woo Ro
Abstract
Recent GPUs leverage Winograd convolution and structured pruning to significantly accelerate inference. First, Winograd convolution is theoretically 2.25× faster than standard convolution. Second, structured pruning reduces inference time without additional overhead as the pruning ratio increases. However, applying conventional structured pruning alongside Winograd convolution is inefficient. Existing structured pruning methods, which do not account for how GPUs process Winograd convolution, require large pruning unit sizes, leading to significant information loss. In this paper, we propose Winograd Structured Pruning (WINS), the first approach to employ optimized structured pruning for Winograd convolution. WINS is designed based on an in-depth analysis of Winograd convolution's computational characteristics on GPUs. Additionally, we introduce two variants, WINS-B and WINS-AB, which further enhance performance. Experimental results show that WINS-AB achieves up to 2.8× practical speedup in baseline inference on GPUs while preserving the accuracy of ResNet-18 on ImageNet.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 80f6e2ea-42f5-4ca1-9a00-6da5ed9e33e0Cited by top-tier papers1
Ask how each one uses itBuilds on18
- Pruning neural networks without any data by iteratively conserving synaptic flowHidenori Tanaka, Daniel Kunin, Daniel L. K. Yamins, Surya GanguliNeurIPS 2020 · 884 citations
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 794 citations
- PatDNN: Achieving Real-Time DNN Execution on Mobile Devices with Pattern-based Weight PruningWei Niu, Xiaolong Ma, Sheng Lin, Shihao Wang et al.ASPLOS 2020 · 214 citations
- Provable Filter Pruning for Efficient Neural NetworksLucas Liebenwein, Cenk Baykal, Harry Lang, Dan Feldman et al.ICLR 2020 · 161 citations
- Structural Pruning via Latency-Saliency KnapsackMaying Shen, Hongxu Yin, Pavlo Molchanov, Lei Mao et al.NeurIPS 2022 · 70 citations
Related papers
- DEPrune: Depth-wise Separable Convolution Pruning for Maximizing GPU ParallelismCheonjun Park, Mincheol Park, Hyunchan Moon, Myung Kuk Yoon et al.NeurIPS 2024 · 10 citations
- WinoTrain: Winograd-Aware Training for Accurate Full 8-bit Convolution AccelerationPierpaolo Morì, Shambhavi Balamuthu Sampath, Lukas Frickenstein, Manoj Rohit Vemparala et al.DAC 2023 · 5 citations
- Dynamic Structure Pruning for Compressing CNNsJun-Hyung Park, Yeachan Kim, Junho Kim, Joon-Young Choi et al.AAAI 2023 · 24 citations
- WinoGen: A Highly Configurable Winograd Convolution IP Generator for Efficient CNN Acceleration on FPGAMingjun Li, Pengjia Li, Shuo Yin, Shixin Chen et al.DAC 2024 · 6 citations
- HBP: Hierarchically Balanced Pruning and Accelerator Co-Design for Efficient DNN InferenceAo Ren, Yuhao Wang, Tao Zhang, Jiaxing Shi et al.DAC 2023 · 4 citations
