Drop an Octave: Reducing Spatial Redundancy in Convolutional Neural Networks With Octave Convolution
Yunpeng Chen, Haoqi Fan, Bing Xu, Zhicheng Yan, Yannis Kalantidis, Marcus Rohrbach, Shuicheng Yan, Jiashi Feng
Abstract
In natural images, information is conveyed at different frequencies where higher frequencies are usually encoded with fine details and lower frequencies are usually encoded with global structures. Similarly, the output feature maps of a convolution layer can also be seen as a mixture of information at different frequencies. In this work, we propose to factorize the mixed feature maps by their frequencies, and design a novel Octave Convolution (OctConv) operation 1 to store and process feature maps that vary spatially "slower" at a lower spatial resolution reducing both memory and computation cost. Unlike existing multi-scale methods, OctConv is formulated as a single, generic, plug-andplay convolutional unit that can be used as a direct replacement of (vanilla) convolutions without any adjustments in the network architecture. It is also orthogonal and complementary to methods that suggest better topologies or reduce channel-wise redundancy like group or depth-wise convolutions. We experimentally show that by simply replacing convolutions with OctConv, we can consistently boost accuracy for both image and video recognition tasks, while reducing memory and computational cost. An OctConv-equipped ResNet-152 can achieve 82.9% top-1 classification accuracy on ImageNet with merely 22.2 GFLOPs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ae129583-b7dc-4177-86d6-3a5b14c53ea2Cited by top-tier papers62
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image ClassificationChun-Fu (Richard) Chen, Quanfu Fan, Rameswar PandaICCV 2021 · 2,072 citations
- Multiscale Vision TransformersHaoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li et al.ICCV 2021 · 1,611 citations
- Fast Fourier ConvolutionLu Chi, Borui Jiang, Yadong MuNeurIPS 2020 · 842 citations
- VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and TextHassan Akbari, Liangzhe Yuan, Rui Qian, Wei-Hong Chuang et al.NeurIPS 2021 · 782 citations
Builds on1
Related papers
- Two-Stage Octave Residual Network for End-to-End Image CompressionFangdong Chen, Yumeng Xu, Li WangAAAI 2022 · 47 citations
- Dual-Octave Convolution for Accelerated Parallel MR Image ReconstructionChun-Mei Feng, Zhanyuan Yang, Geng Chen, Yong Xu et al.AAAI 2021 · 31 citations
- Diverse Branch Block: Building a Convolution as an Inception-Like UnitXiaohan Ding, Xiangyu Zhang, Jungong Han, Guiguang DingCVPR 2021
- Learned Bi-Resolution Image Coding using Generalized Octave ConvolutionsMohammad Akbari, Jie Liang, Jingning Han, Chengjie TuAAAI 2021 · 21 citations
- Learning in the Frequency DomainKai Xu, Minghai Qin, Fei Sun, Yuhao Wang et al.CVPR 2020
