Poly-NL: Linear Complexity Non-local Layers With 3rd Order Polynomials
Francesca Babiloni, Ioannis Marras, Filippos Kokkinos, Jiankang Deng, Grigorios Chrysos, Stefanos Zafeiriou
Abstract
Spatial self-attention layers, in the form of Non-Local blocks, introduce long-range dependencies in Convolutional Neural Networks by computing pairwise similarities among all possible positions. Such pairwise functions underpin the effectiveness of non-local layers, but also determine a complexity that scales quadratically with respect to the input size both in space and time. This is a severely limiting factor that practically hinders the applicability of non-local blocks to even moderately sized inputs. Previous works focused on reducing the complexity by modifying the underlying matrix operations, however in this work we aim to retain full expressiveness of non-local layers while keeping complexity linear. We overcome the efficiency limitation of non-local blocks by framing them as special cases of 3rd order polynomial functions. This fact enables us to formulate novel fast Non-Local blocks, capable of reducing the complexity from quadratic to linear with no loss in performance, by replacing any direct computation of pairwise similarities with element-wise multiplications. The proposed method, which we dub as "Poly-NL", is competitive with state-of-the-art performance across image recognition, instance segmentation, and face detection tasks, while having considerably less computational overhead.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- MB-TaylorFormer: Multi-branch Efficient Transformer Expanded by Taylor Formula for Image DehazingYuwei Qiu, Kaihao Zhang, Chenxi Wang, Wenhan Luo et al.ICCV 2023 · 224 citations
- Multilinear Mixture of Experts: Scalable Expert Specialization through FactorizationJames Oldfield, Markos Georgopoulos, Grigorios Chrysos, Christos Tzelepis et al.NeurIPS 2024 · 41 citations
- Extrapolation and Spectral Bias of Neural Nets with Hadamard Product: a Polynomial Net StudyYongtao Wu, Zhenyu Zhu, Fanghui Liu, Grigorios Chrysos et al.NeurIPS 2022 · 19 citations
- QT-ViT: Improving Linear Attention in ViT with Quadratic Taylor ExpansionYixing Xu, Chao Li, Dong Li, Xiao Sheng et al.NeurIPS 2024 · 7 citations
- PolySAE: Modeling Feature Interactions in Sparse Autoencoders via Polynomial DecodingPanagiotis Koromilas, Andreas Demou, James Oldfield, Yannis Panagakis et al.ICML 2026 · 3 citations
Builds on7
- Distance-IoU Loss: Faster and Better Learning for Bounding Box RegressionZhaohui Zheng, Ping Wang, Wei Liu, Jinze Li et al.AAAI 2020 · 4,823 citations
- Attention Augmented Convolutional NetworksIrwan Bello, Barret Zoph, Quoc Le, Ashish Vaswani et al.ICCV 2019 · 1,149 citations
- Multiplicative Interactions and Where to Find ThemSiddhant M. Jayakumar, Wojciech M. Czarnecki, Jacob Menick, Jonathan Schwarz et al.ICLR 2020 · 152 citations
- Rethinking Attention with PerformersKrzysztof Marcin Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song et al.ICLR 2021 · 122 citations
- TESA: Tensor Element Self-Attention via MatricizationFrancesca Babiloni, Ioannis Marras, Gregory G. Slabaugh, Stefanos ZafeiriouCVPR 2020
Related papers
- Global Filter Networks for Image ClassificationYongming Rao, Wenliang Zhao, Zheng Zhu, Jiwen Lu et al.NeurIPS 2021 · 798 citations
- Extremely Compact Non-local Representation LearningAnsheng You, Xiangzeng Zhou, Yingya Zhang, Pan Pan et al.KDD 2021
- Unifying Nonlocal Blocks for Neural NetworksLei Zhu, Qi She, Duo Li, Yanye Lu et al.ICCV 2021 · 26 citations
- On the Relationship between Self-Attention and Convolutional LayersJean-Baptiste Cordonnier, Andreas Loukas, Martin JaggiICLR 2020 · 629 citations
- LambdaNetworks: Modeling long-range Interactions without AttentionIrwan BelloICLR 2021 · 48 citations
