More ConvNets in the 2020s: Scaling up Kernels Beyond 51x51 using Sparsity
Shiwei Liu, Tianlong Chen, Xiaohan Chen, Xuxi Chen, Qiao Xiao, Boqian Wu, Tommi Kärkkäinen, Mykola Pechenizkiy, Decebal Constantin Mocanu, Zhangyang Wang
摘要
Transformers have quickly shined in the computer vision world since the emergence of Vision Transformers (ViTs). The dominant role of convolutional neural networks (CNNs) seems to be challenged by increasingly effective transformer-based models. Very recently, a couple of advanced convolutional models strike back with large kernels motivated by the local-window attention mechanism, showing appealing performance and efficiency. While one of them, i.e. RepLKNet, impressively manages to scale the kernel size to 31×31 with improved performance, the performance starts to saturate as the kernel size continues growing, compared to the scaling trend of advanced ViTs such as Swin Transformer. In this paper, we explore the possibility of training extreme convolutions larger than 31×31 and test whether the performance gap can be eliminated by strategically enlarging convolutions. This study ends up with a recipe for applying extremely large kernels from the perspective of sparsity, which can smoothly scale up kernels to 61×61 with better performance. Built on this recipe, we propose Sparse Large Kernel Network (SLaK), a pure CNN architecture equipped with sparse factorized 51×51 kernels that can perform on par with or better than state-of-the-art hierarchical Transformers and modern ConvNet architectures like ConvNeXt and RepLKNet, on ImageNet classification as well as a wide range of downstream tasks including semantic segmentation on ADE20K, object detection on PASCAL VOC 2007, and object detection/segmentation on MS COCO. Codes: https://github.com/VITA-Group/SLaK * Equal contribution.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper69
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu 等NeurIPS 2024 · 被引用 3,199 次
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang 等ICML 2024 · 被引用 1,725 次
- Large Selective Kernel Network for Remote Sensing Object DetectionYuxuan Li, Qibin Hou, Zhaohui Zheng, Ming-Ming Cheng 等ICCV 2023 · 被引用 535 次
- ModernTCN: A Modern Pure Convolution Structure for General Time Series AnalysisDonghao Luo, Xue WangICLR 2024 · 被引用 392 次
- InceptionNeXt: When Inception Meets ConvNeXtWeihao Yu, Pan Zhou, Shuicheng Yan, Xinchao WangCVPR 2024 · 被引用 326 次
它引用的顶会 Paper42
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
相关 Paper
- Scaling Up Your Kernels to 31×31: Revisiting Large Kernel Design in CNNsXiaohan Ding, Xiangyu Zhang, Jungong Han, Guiguang DingCVPR 2022 · 被引用 1,298 次
- CAM Back Again: Large Kernel CNNs from a Weakly Supervised Object Localization PerspectiveShunsuke Yasuki, Masato TakiCVPR 2024 · 被引用 6 次
- ShiftwiseConv: Small Convolutional Kernel with Large Kernel EffectDachong Li, Li Li, Zhuangzhuang Chen, Jianqiang LiCVPR 2025
- PeLK: Parameter-Efficient Large Kernel ConvNets with Peripheral ConvolutionHonghao Chen, Xiangxiang Chu, Yongjian Ren, Xin Zhao 等CVPR 2024
- Are Large Kernels Better Teachers than Transformers for ConvNets?Tianjin Huang, Lu Yin, Zhenyu Zhang, Li Shen 等ICML 2023 · 被引用 18 次
