Non-Local Neural Networks With Grouped Bilinear Attentional Transforms
Lu Chi, Zehuan Yuan, Yadong Mu, Changhu Wang
Abstract
Modeling spatial or temporal long-range dependency plays a key role in deep neural networks. Conventional dominant solutions include recurrent operations on sequential data or deeply stacking convolutional layers with small kernel size. Recently, a number of non-local operators (such as self-attention based [57]) have been devised. They are typically generic and can be plugged into many existing network pipelines for globally computing among any two neurons in a feature map. This work proposes a novel non-local operator. It is inspired by the attention mechanism of human visual system, which can quickly attend to important local parts in sight and suppress other less-relevant information. The core of our method is learnable and data-adaptive bilinear attentional transform (BA-Transform), whose merits are three-folds: first, BA-Transform is versatile to model a wide spectrum of local or global attentional operations, such as emphasizing specific local regions. Each BA-Transform is learned in a dataadaptive way; Secondly, to address the discrepancy among features, we further design grouped BA-Transforms, which essentially apply different attentional operations to different groups of feature channels; Thirdly, many existing nonlocal operators are computation-intensive. The proposed BA-Transform is implemented by simple matrix multiplication and admits better efficacy. For empirical evaluation, we perform comprehensive experiments on two large-scale benchmarks, ImageNet and Kinetics, for image / video classification respectively. The achieved accuracies and various ablation experiments consistently demonstrate significant improvement by large margins.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Dual Contrastive Loss and Attention for GANsNing Yu, Guilin Liu, Aysegul Dundar, Andrew Tao et al.ICCV 2021 · 69 citations
- Temporal-attentive Covariance Pooling Networks for Video RecognitionZilin Gao, Qilong Wang, Bingbing Zhang, Qinghua Hu et al.NeurIPS 2021 · 33 citations
- FFNet: Frequency Fusion Network for Semantic Scene CompletionXuzhi Wang, Di Lin, Liang WanAAAI 2022 · 28 citations
Builds on8
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le et al.ICCV 2019 · 9,163 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- CCNet: Criss-Cross Attention for Semantic SegmentationZilong Huang, Xinggang Wang, Lichao Huang, Chang Huang et al.ICCV 2019 · 2,972 citations
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- Progressive Differentiable Architecture Search: Bridging the Depth Gap Between Search and EvaluationXin Chen, Lingxi Xie, Jun Wu, Qi TianICCV 2019 · 725 citations
Related papers
- Unifying Nonlocal Blocks for Neural NetworksLei Zhu, Qi She, Duo Li, Yanye Lu et al.ICCV 2021 · 26 citations
- Relational Self-Attention: What's Missing in Attention for Video UnderstandingManjin Kim, Heeseung Kwon, Chunyu Wang, Suha Kwak et al.NeurIPS 2021 · 40 citations
- Fast Fourier ConvolutionLu Chi, Borui Jiang, Yadong MuNeurIPS 2020 · 842 citations
- KNN Local Attention for Image RestorationHunsang Lee, Hyesong Choi, Kwanghoon Sohn, Dongbo MinCVPR 2022 · 62 citations
- UniFormer: Unified Transformer for Efficient Spatial-Temporal Representation LearningKunchang Li, Yali Wang, Peng Gao, Guanglu Song et al.ICLR 2022
