Non-Local Neural Networks With Grouped Bilinear Attentional Transforms
Lu Chi, Zehuan Yuan, Yadong Mu, Changhu Wang
摘要
Modeling spatial or temporal long-range dependency plays a key role in deep neural networks. Conventional dominant solutions include recurrent operations on sequential data or deeply stacking convolutional layers with small kernel size. Recently, a number of non-local operators (such as self-attention based [57]) have been devised. They are typically generic and can be plugged into many existing network pipelines for globally computing among any two neurons in a feature map. This work proposes a novel non-local operator. It is inspired by the attention mechanism of human visual system, which can quickly attend to important local parts in sight and suppress other less-relevant information. The core of our method is learnable and data-adaptive bilinear attentional transform (BA-Transform), whose merits are three-folds: first, BA-Transform is versatile to model a wide spectrum of local or global attentional operations, such as emphasizing specific local regions. Each BA-Transform is learned in a dataadaptive way; Secondly, to address the discrepancy among features, we further design grouped BA-Transforms, which essentially apply different attentional operations to different groups of feature channels; Thirdly, many existing nonlocal operators are computation-intensive. The proposed BA-Transform is implemented by simple matrix multiplication and admits better efficacy. For empirical evaluation, we perform comprehensive experiments on two large-scale benchmarks, ImageNet and Kinetics, for image / video classification respectively. The achieved accuracies and various ablation experiments consistently demonstrate significant improvement by large margins.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Dual Contrastive Loss and Attention for GANsNing Yu, Guilin Liu, Aysegul Dundar, Andrew Tao 等ICCV 2021 · 被引用 69 次
- Temporal-attentive Covariance Pooling Networks for Video RecognitionZilin Gao, Qilong Wang, Bingbing Zhang, Qinghua Hu 等NeurIPS 2021 · 被引用 33 次
- FFNet: Frequency Fusion Network for Semantic Scene CompletionXuzhi Wang, Di Lin, Liang WanAAAI 2022 · 被引用 28 次
它引用的顶会 Paper8
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le 等ICCV 2019 · 被引用 9,163 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- CCNet: Criss-Cross Attention for Semantic SegmentationZilong Huang, Xinggang Wang, Lichao Huang, Chang Huang 等ICCV 2019 · 被引用 2,972 次
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 被引用 2,049 次
- Progressive Differentiable Architecture Search: Bridging the Depth Gap Between Search and EvaluationXin Chen, Lingxi Xie, Jun Wu, Qi TianICCV 2019 · 被引用 725 次
相关 Paper
- Unifying Nonlocal Blocks for Neural NetworksLei Zhu, Qi She, Duo Li, Yanye Lu 等ICCV 2021 · 被引用 26 次
- Relational Self-Attention: What's Missing in Attention for Video UnderstandingManjin Kim, Heeseung Kwon, Chunyu Wang, Suha Kwak 等NeurIPS 2021 · 被引用 40 次
- Fast Fourier ConvolutionLu Chi, Borui Jiang, Yadong MuNeurIPS 2020 · 被引用 842 次
- KNN Local Attention for Image RestorationHunsang Lee, Hyesong Choi, Kwanghoon Sohn, Dongbo MinCVPR 2022 · 被引用 62 次
- UniFormer: Unified Transformer for Efficient Spatial-Temporal Representation LearningKunchang Li, Yali Wang, Peng Gao, Guanglu Song 等ICLR 2022
