Involution: Inverting the Inherence of Convolution for Visual Recognition
Duo Li, Jie Hu, Changhu Wang, Xiangtai Li, Qi She, Lei Zhu, Tong Zhang, Qifeng Chen
摘要
Convolution has been the core ingredient of modern neural networks, triggering the surge of deep learning in vision. In this work, we rethink the inherent principles of standard convolution for vision tasks, specifically spatialagnostic and channel-specific. Instead, we present a novel atomic operation for deep neural networks by inverting the aforementioned design principles of convolution, coined as involution. We additionally demystify the recent popular self-attention operator and subsume it into our involution family as an over-complicated instantiation. The proposed involution operator could be leveraged as fundamental bricks to build the new generation of neural networks for visual recognition, powering different deep learning models on several prevalent benchmarks, including Im-ageNet classification, COCO detection and segmentation, together with Cityscapes segmentation. Our involutionbased models improve the performance of convolutional baselines using ResNet-50 by up to 1.6% top-1 accuracy, 2.5% and 2.4% bounding box AP, and 4.7% mean IoU absolutely while compressing the computational cost to 66%, 65%, 72%, and 57% on the above benchmarks, respectively. Code and pre-trained models for all the tasks are available at https://github.com/d-li14/involution .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper34
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer 等NeurIPS 2021 · 被引用 3,862 次
- Focal Modulation NetworksJianwei Yang, Chunyuan Li, Xiyang Dai, Jianfeng GaoNeurIPS 2022 · 被引用 494 次
- TransNeXt: Robust Foveal Visual Perception for Vision TransformersDai ShiCVPR 2024 · 被引用 313 次
- On the Connection between Local Attention and Dynamic Depth-wise ConvolutionQi Han, Zejia Fan, Qi Dai, Lei Sun 等ICLR 2022 · 被引用 144 次
- MixFormer: Mixing Features across Windows and DimensionsQiang Chen, Qiman Wu, Jian Wang, Qinghao Hu 等CVPR 2022 · 被引用 142 次
它引用的顶会 Paper13
- CCNet: Criss-Cross Attention for Semantic SegmentationZilong Huang, Xinggang Wang, Lichao Huang, Chang Huang 等ICCV 2019 · 被引用 2,972 次
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy 等ICCV 2019 · 被引用 1,396 次
- Attention Augmented Convolutional NetworksIrwan Bello, Barret Zoph, Quoc Le, Ashish Vaswani 等ICCV 2019 · 被引用 1,149 次
- CARAFE: Content-Aware ReAssembly of FEaturesJiaqi Wang, Kai Chen, Rui Xu, Ziwei Liu 等ICCV 2019 · 被引用 842 次
- On the Relationship between Self-Attention and Convolutional LayersJean-Baptiste Cordonnier, Andreas Loukas, Martin JaggiICLR 2020 · 被引用 629 次
相关 Paper
- MOAT: Alternating Mobile Convolution and Attention Brings Strong Vision ModelsChenglin Yang, Siyuan Qiao, Qihang Yu, Xiaoding Yuan 等ICLR 2023 · 被引用 22 次
- Bottleneck Transformers for Visual RecognitionAravind Srinivas, Tsung-Yi Lin, Niki Parmar, Jonathon Shlens 等CVPR 2021
- PartialNet: Compute Less, Perform BetterHaiduo Huang, Tian Xia, Wenzhe Zhao, Pengju RenAAAI 2026
- Reversible Column NetworksYuxuan Cai, Yizhuang Zhou, Qi Han, Jianjian Sun 等ICLR 2023 · 被引用 21 次
- Scaling Local Self-Attention for Parameter Efficient Visual BackbonesAshish Vaswani, Prajit Ramachandran, Aravind Srinivas, Niki Parmar 等CVPR 2021
