Drop an Octave: Reducing Spatial Redundancy in Convolutional Neural Networks With Octave Convolution
Yunpeng Chen, Haoqi Fan, Bing Xu, Zhicheng Yan, Yannis Kalantidis, Marcus Rohrbach, Shuicheng Yan, Jiashi Feng
摘要
In natural images, information is conveyed at different frequencies where higher frequencies are usually encoded with fine details and lower frequencies are usually encoded with global structures. Similarly, the output feature maps of a convolution layer can also be seen as a mixture of information at different frequencies. In this work, we propose to factorize the mixed feature maps by their frequencies, and design a novel Octave Convolution (OctConv) operation 1 to store and process feature maps that vary spatially "slower" at a lower spatial resolution reducing both memory and computation cost. Unlike existing multi-scale methods, OctConv is formulated as a single, generic, plug-andplay convolutional unit that can be used as a direct replacement of (vanilla) convolutions without any adjustments in the network architecture. It is also orthogonal and complementary to methods that suggest better topologies or reduce channel-wise redundancy like group or depth-wise convolutions. We experimentally show that by simply replacing convolutions with OctConv, we can consistently boost accuracy for both image and video recognition tasks, while reducing memory and computational cost. An OctConv-equipped ResNet-152 can achieve 82.9% top-1 classification accuracy on ImageNet with merely 22.2 GFLOPs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper62
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 被引用 2,927 次
- CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image ClassificationChun-Fu (Richard) Chen, Quanfu Fan, Rameswar PandaICCV 2021 · 被引用 2,072 次
- Multiscale Vision TransformersHaoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li 等ICCV 2021 · 被引用 1,611 次
- Fast Fourier ConvolutionLu Chi, Borui Jiang, Yadong MuNeurIPS 2020 · 被引用 842 次
- VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and TextHassan Akbari, Liangzhe Yuan, Rui Qian, Wei-Hong Chuang 等NeurIPS 2021 · 被引用 782 次
它引用的顶会 Paper1
相关 Paper
- Two-Stage Octave Residual Network for End-to-End Image CompressionFangdong Chen, Yumeng Xu, Li WangAAAI 2022 · 被引用 47 次
- Dual-Octave Convolution for Accelerated Parallel MR Image ReconstructionChun-Mei Feng, Zhanyuan Yang, Geng Chen, Yong Xu 等AAAI 2021 · 被引用 31 次
- Diverse Branch Block: Building a Convolution as an Inception-Like UnitXiaohan Ding, Xiangyu Zhang, Jungong Han, Guiguang DingCVPR 2021
- Learned Bi-Resolution Image Coding using Generalized Octave ConvolutionsMohammad Akbari, Jie Liang, Jingning Han, Chengjie TuAAAI 2021 · 被引用 21 次
- Learning in the Frequency DomainKai Xu, Minghai Qin, Fei Sun, Yuhao Wang 等CVPR 2020
