Fast Fourier Convolution
Lu Chi, Borui Jiang, Yadong Mu
Abstract
Vanilla convolutions in modern deep networks are known to operate locally and at fixed scale (e.g., the widely-adopted 3 × 3 kernels in image-oriented tasks). This causes low efficacy in connecting two distant locations in the network. In this work, we propose a novel convolutional operator dubbed as fast Fourier convolution (FFC), which has the main hallmarks of non-local receptive fields and cross-scale fusion within the convolutional unit. According to spectral convolution theorem in Fourier theory, point-wise update in the spectral domain globally affects all input features involved in Fourier transform, which sheds light on neural architectural design with non-local receptive field. Our proposed FFC is inspired to capsulate three different kinds of computations in a single operation unit: a local branch that conducts ordinary small-kernel convolution, a semi-global branch that processes spectrally stacked image patches, and a global branch that manipulates image-level spectrum. All branches complementarily address different scales. A multi-branch aggregation step is included in FFC for cross-scale fusion. FFC is a generic operator that can directly replace vanilla convolutions in a large body of existing networks, without any adjustments and with comparable complexity metrics (e.g., FLOPs). We experimentally evaluate FFC in three major vision benchmarks (ImageNet for image recognition, Kinetics for video action recognition, MSCOCO for human keypoint detection). It consistently elevates accuracies in all above tasks by significant margins.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 27441411-3ba2-4b59-9e8c-b3efea181c8fCited by top-tier papers65
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Fourier Neural Operator for Parametric Partial Differential EquationsZongyi Li, Nikola Borislavov Kovachki, Kamyar Azizzadenesheli, Burigede Liu et al.ICLR 2021 · 3,911 citations
- Incremental Transformer Structure Enhanced Image Inpainting with Masking Positional EncodingQiaole Dong, Chenjie Cao, Yanwei FuCVPR 2022 · 194 citations
- Fourmer: An Efficient Global Modeling Paradigm for Image RestorationMan Zhou, Jie Huang, Chun-Le Guo, Chongyi LiICML 2023 · 148 citations
- Frequency Enhanced Hybrid Attention Network for Sequential RecommendationXinyu Du, Huanhuan Yuan, Pengpeng Zhao, Jianfeng Qu et al.SIGIR 2023 · 142 citations
Builds on3
- CCNet: Criss-Cross Attention for Semantic SegmentationZilong Huang, Xinggang Wang, Lichao Huang, Chang Huang et al.ICCV 2019 · 2,972 citations
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- Drop an Octave: Reducing Spatial Redundancy in Convolutional Neural Networks With Octave ConvolutionYunpeng Chen, Haoqi Fan, Bing Xu, Zhicheng Yan et al.ICCV 2019 · 665 citations
Related papers
- Deep Fourier Up-SamplingMan Zhou, Hu Yu, Jie Huang, Feng Zhao et al.NeurIPS 2022 · 80 citations
- Adaptive Frequency Filters As Efficient Global Token MixersZhipeng Huang, Zhizheng Zhang, Cuiling Lan, Zheng-Jun Zha et al.ICCV 2023 · 96 citations
- Frequency-Adaptive Dilated Convolution for Semantic SegmentationLinwei Chen, Lin Gu, Dezhi Zheng, Ying FuCVPR 2024
- UniConvNet: Expanding Effective Receptive Field While Maintaining Asymptotically Gaussian Distribution for ConvNets of Any ScaleYuhao Wang, Wei XiICCV 2025 · 5 citations
- Interferometric Graph Transform: a Deep Unsupervised Graph RepresentationEdouard OyallonICML 2020 · 6 citations
