Spatial Assembly Networks for Image Representation Learning
Yang Li, Shichao Kan, Jianhe Yuan, Wenming Cao, Zhihai He
Abstract
It has been long recognized that deep neural networks are sensitive to changes in spatial configurations or scene structures. Image augmentations, such as random translation, cropping, and resizing, can be used to improve the robustness of deep neural networks under spatial transforms. However, changes in object part configurations, spatial layout of object, and scene structures of the images may still result in major changes in the their feature representations generated by the network, creating significant challenges for various visual learning tasks, including representation or metric learning, image classification and retrieval. In this work, we introduce a new learnable module, called spatial assembly network (SAN), to address this important issue. This SAN module examines the input image and performs a learned re-organization and assembly of feature points from different spatial locations conditioned by feature maps from previous network layers so as to maximize the discriminative power of the final feature representation. This differentiable module can be flexibly incorporated into existing network architectures, improving their capabilities in handling spatial variations and structural changes of the image scene. We demonstrate that the proposed SAN module is able to significantly improve the performance of various metric / representation learning, image retrieval and classification tasks, in both supervised and unsupervised learning scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on6
- AutoShuffleNet: Learning Permutation Matrices via an Exact Lipschitz Continuous Penalty in Deep Convolutional Neural NetworksJiancheng Lyu, Shuai Zhang, Yingyong Qi, Jack XinKDD 2020 · 18 citations
- Probabilistic Structural Latent Representation for Unsupervised EmbeddingMang Ye, Jianbing ShenCVPR 2020
- Cross-Batch Memory for Embedding LearningXun Wang, Haozhi Zhang, Weilin Huang, Matthew R. ScottCVPR 2020
- On Translation Invariance in CNNs: Convolutional Layers Can Exploit Absolute Spatial LocationOsman Semih Kayhan, Jan C. van GemertCVPR 2020
- Proxy Anchor Loss for Deep Metric LearningSungyeon Kim, Dongwon Kim, Minsu Cho, Suha KwakCVPR 2020
Related papers
- Rethinking Alignment and Uniformity in Unsupervised Image Semantic SegmentationDaoan Zhang, Chenming Li, Haoquan Li, Wenjian Huang et al.AAAI 2023 · 21 citations
- Revisiting Self-Similarity: Structural Embedding for Image RetrievalSeongwon Lee, Suhyeon Lee, Hongje Seong, Euntai KimCVPR 2023
- Learning Spatial-context-aware Global Visual Feature Representation for Instance Image RetrievalZhongyan Zhang, Lei Wang, Luping Zhou, Piotr KoniuszICCV 2023 · 13 citations
- Stochastic Attraction-Repulsion Embedding for Large Scale Image LocalizationLiu Liu, Hongdong Li, Yuchao DaiICCV 2019 · 123 citations
- Spatial-Aware Token for Weakly Supervised Object LocalizationPingyu Wu, Wei Zhai, Yang Cao, Jiebo Luo et al.ICCV 2023 · 19 citations
