Learning Spatial-context-aware Global Visual Feature Representation for Instance Image Retrieval
Zhongyan Zhang, Lei Wang, Luping Zhou, Piotr Koniusz
Abstract
In instance image retrieval, considering local spatial information within an image has proven effective to boost retrieval performance, as demonstrated by local visual descriptor based geometric verification. Nevertheless, it will be highly valuable to make ordinary global image representations spatial-context-aware because global representation based image retrieval is appealing thanks to its algorithmic simplicity, low memory cost, and being friendly to sophisticated data structures. To this end, we propose a novel feature learning framework for instance image retrieval, which embeds local spatial context information into the learned global feature representations. Specifically, in parallel to the visual feature branch in a CNN backbone, we design a spatial context branch that consists of two modules called online token learning and distance encoding. For each local descriptor learned in CNN, the former module is used to indicate the types of its surrounding descriptors, while their spatial distribution information is captured by the latter module. After that, the visual feature branch and the spatial context branch are fused to produce a single global feature representation per image. As experimentally demonstrated, with the spatial-context-aware characteristic, we can well improve the performance of global representation based image retrieval while maintaining all of its appealing properties. Our code is available at https://github.com/Zy-Zhang/SpCa.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 493140d1-b40d-44da-9f1b-b6d7f425008bCited by top-tier papers2
- Image Clustering Conditioned on Text CriteriaSehyun Kwon, Jaeseung Park, Minkyu Kim, Jaewoong Cho et al.ICLR 2024 · 27 citations
- Robust Distillation via Untargeted and Targeted Intermediate Adversarial SamplesJunhao Dong, Piotr Koniusz, Junxi Chen, Z. Jane Wang et al.CVPR 2024
Builds on11
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Learning With Average Precision: Training Image Retrieval With a Listwise LossJérôme Revaud, Jon Almazán, Rafael S. Rezende, César Roberto de SouzaICCV 2019 · 424 citations
- Distance Encoding: Design Provably More Powerful Neural Networks for Graph Representation LearningPan Li, Yanbang Wang, Hongwei Wang, Jure LeskovecNeurIPS 2020 · 391 citations
- DOLG: Single-Stage Image Retrieval with Deep Orthogonal Fusion of Local and Global FeaturesMin Yang, Dongliang He, Miao Fan, Baorong Shi et al.ICCV 2021 · 135 citations
- Instance-level Image Retrieval using Reranking TransformersFuwen Tan, Jiangbo Yuan, Vicente OrdonezICCV 2021 · 116 citations
Related papers
- Learning Token-Based Representation for Image RetrievalHui Wu, Min Wang, Wengang Zhou, Yang Hu et al.AAAI 2022 · 26 citations
- StructVPR: Distill Structural Knowledge with Weighting Samples for Visual Place RecognitionYanqing Shen, Sanping Zhou, Jingwen Fu, Ruotong Wang et al.CVPR 2023
- SCORE: Scene Context Matters in Open-Vocabulary Remote Sensing Instance SegmentationShiqi Huang, Shuting He, Huaiyuan Qin, Bihan WenICCV 2025 · 2 citations
- Improving Fusion of Region Features and Grid Features via Two-Step Interaction for Image-Text RetrievalDongqing Wu, Huihui Li, Cang Gu, Lei Guo et al.ACM MM 2022 · 10 citations
- Revisiting Self-Similarity: Structural Embedding for Image RetrievalSeongwon Lee, Suhyeon Lee, Hongje Seong, Euntai KimCVPR 2023
