TESA: Tensor Element Self-Attention via Matricization
Francesca Babiloni, Ioannis Marras, Gregory G. Slabaugh, Stefanos Zafeiriou
Abstract
Representation learning is a fundamental part of modern computer vision, where abstract representations of data are encoded as tensors optimized to solve problems like image segmentation and inpainting. Recently, self-attention in the form of a Non-Local Block has emerged as a powerful technique to enrich features, by capturing complex interdependencies in feature tensors. However, standard selfattention approaches leverage only spatial relationships, drawing similarities between vectors and overlooking correlations between channels. In this paper, we introduce a new method, called Tensor Element Self-Attention (TESA) that generalizes such work to capture interdependencies along all dimensions of the tensor using matricization. An order R tensor produces R results, one for each dimension. The results are then fused to produce an enriched output which encapsulates similarity among tensor elements. Additionally, we analyze self-attention mathematically, providing new perspectives on how it adjusts the singular values of the input feature tensor. With these new insights, we present experimental results demonstrating how TESA can benefit diverse problems including classification and instance segmentation. By simply adding a TESA module to existing networks, we substantially improve competitive baselines and set new state-of-the-art results for image inpainting on CelebA and low light raw-to-rgb image translation on SID.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Is Attention Better Than Matrix Decomposition?Zhengyang Geng, Meng-Hao Guo, Hongxu Chen, Xia Li et al.ICLR 2021 · 171 citations
- Fully Attentional Network for Semantic SegmentationQi Song, Jie Li, Chenghong Li, Hao Guo et al.AAAI 2022 · 63 citations
- Multilinear Mixture of Experts: Scalable Expert Specialization through FactorizationJames Oldfield, Markos Georgopoulos, Grigorios Chrysos, Christos Tzelepis et al.NeurIPS 2024 · 41 citations
- Poly-NL: Linear Complexity Non-local Layers With 3rd Order PolynomialsFrancesca Babiloni, Ioannis Marras, Filippos Kokkinos, Jiankang Deng et al.ICCV 2021 · 14 citations
Builds on2
Related papers
- Steering Self-Supervised Feature Learning Beyond Local Pixel StatisticsSimon Jenni, Hailin Jin, Paolo FavaroCVPR 2020
- Attentive Normalization for Conditional Image GenerationYi Wang, Ying-Cong Chen, Xiangyu Zhang, Jian Sun et al.CVPR 2020
- UCTGAN: Diverse Image Inpainting Based on Unsupervised Cross-Space TranslationLei Zhao, Qihang Mo, Sihuan Lin, Zhizhong Wang et al.CVPR 2020
- Unifying Nonlocal Blocks for Neural NetworksLei Zhu, Qi She, Duo Li, Yanye Lu et al.ICCV 2021 · 26 citations
- TS-CAM: Token Semantic Coupled Attention Map for Weakly Supervised Object LocalizationWei Gao, Fang Wan, Xingjia Pan, Zhiliang Peng et al.ICCV 2021 · 260 citations
