WBMM: Windowed Batch Matrix Multiplication for Efficient Large Receptive Field Convolution
Wan Song, Zhou Wei, Rui Wang, Jun Yu, Toru Kurihara, Xu Jiajia, shu zhan
Abstract
Large kernel depthwise convolutions achieve strong performance but suffer from significant degradation as kernel size grows due to irregular memory access from gather-based computation. While Large Kernel Acceleration (LKA) helps on small feature maps, it becomes counterproductive on large feature maps, even slower than non-accelerated implementations. We propose Windowed Batch Matrix Multiplication (WBMM), which partitions input into contiguous windows and indexes a compact relative position bias table to construct weight matrices, enabling regular memory access via batched matrix multiplication; this yields a unique property where WBMM's throughput improves with larger windows, opposite to depthwise convolutions that degrade with larger kernels. Operator-level benchmarks show WBMM with windows outperforms depthwise convolution baselines in speed while providing larger receptive field, and combined with inter-block cross-window communication and hierarchical window reparameterization, achieves comparable or higher accuracy on ImageNet-1K, COCO, and ADE20K with 1.31--1.88 training speedup. WBMM also demonstrates consistent advantages across diverse hardware platforms including GPU, CPU, and edge devices, without requiring specialized acceleration kernels. Code and models will be publicly available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on9
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- ACNet: Strengthening the Kernel Skeletons for Powerful CNN via Asymmetric Convolution BlocksXiaohan Ding, Yuchen Guo, Guiguang Ding, Jungong HanICCV 2019 · 845 citations
- Conditional Positional Encodings for Vision TransformersXiangxiang Chu, Zhi Tian, Bo Zhang, Xinlong Wang et al.ICLR 2023 · 406 citations
Related papers
- InceptionNeXt: When Inception Meets ConvNeXtWeihao Yu, Pan Zhou, Shuicheng Yan, Xinchao WangCVPR 2024 · 326 citations
- DECA: A Near-Core LLM Decompression Accelerator Grounded on a 3D Roofline ModelGerasimos Gerogiannis, Stijn Eyerman, Evangelos Georganas, Wim Heirman et al.MICRO 2025 · 5 citations
- Spada: Accelerating Sparse Matrix Multiplication with Adaptive DataflowZhiyao Li, Jiaxiang Li, Taijie Chen, Dimin Niu et al.ASPLOS 2023 · 59 citations
- DWM: A Decomposable Winograd Method for Convolution AccelerationDi Huang, Xishan Zhang, Rui Zhang, Tian Zhi et al.AAAI 2020 · 31 citations
- BlockPIM: Optimizing Memory Management for PIM-enabled Long-Context LLM InferenceZhichun Li, Jun Zhou, Xueqi Li, Ninghui SunDAC 2025 · 3 citations
