Learning Spatial-context-aware Global Visual Feature Representation for Instance Image Retrieval
Zhongyan Zhang, Lei Wang, Luping Zhou, Piotr Koniusz
摘要
In instance image retrieval, considering local spatial information within an image has proven effective to boost retrieval performance, as demonstrated by local visual descriptor based geometric verification. Nevertheless, it will be highly valuable to make ordinary global image representations spatial-context-aware because global representation based image retrieval is appealing thanks to its algorithmic simplicity, low memory cost, and being friendly to sophisticated data structures. To this end, we propose a novel feature learning framework for instance image retrieval, which embeds local spatial context information into the learned global feature representations. Specifically, in parallel to the visual feature branch in a CNN backbone, we design a spatial context branch that consists of two modules called online token learning and distance encoding. For each local descriptor learned in CNN, the former module is used to indicate the types of its surrounding descriptors, while their spatial distribution information is captured by the latter module. After that, the visual feature branch and the spatial context branch are fused to produce a single global feature representation per image. As experimentally demonstrated, with the spatial-context-aware characteristic, we can well improve the performance of global representation based image retrieval while maintaining all of its appealing properties. Our code is available at https://github.com/Zy-Zhang/SpCa.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Image Clustering Conditioned on Text CriteriaSehyun Kwon, Jaeseung Park, Minkyu Kim, Jaewoong Cho 等ICLR 2024 · 被引用 27 次
- Robust Distillation via Untargeted and Targeted Intermediate Adversarial SamplesJunhao Dong, Piotr Koniusz, Junxi Chen, Z. Jane Wang 等CVPR 2024
它引用的顶会 Paper11
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Learning With Average Precision: Training Image Retrieval With a Listwise LossJérôme Revaud, Jon Almazán, Rafael S. Rezende, César Roberto de SouzaICCV 2019 · 被引用 424 次
- Distance Encoding: Design Provably More Powerful Neural Networks for Graph Representation LearningPan Li, Yanbang Wang, Hongwei Wang, Jure LeskovecNeurIPS 2020 · 被引用 391 次
- DOLG: Single-Stage Image Retrieval with Deep Orthogonal Fusion of Local and Global FeaturesMin Yang, Dongliang He, Miao Fan, Baorong Shi 等ICCV 2021 · 被引用 135 次
- Instance-level Image Retrieval using Reranking TransformersFuwen Tan, Jiangbo Yuan, Vicente OrdonezICCV 2021 · 被引用 116 次
相关 Paper
- Learning Token-Based Representation for Image RetrievalHui Wu, Min Wang, Wengang Zhou, Yang Hu 等AAAI 2022 · 被引用 26 次
- StructVPR: Distill Structural Knowledge with Weighting Samples for Visual Place RecognitionYanqing Shen, Sanping Zhou, Jingwen Fu, Ruotong Wang 等CVPR 2023
- SCORE: Scene Context Matters in Open-Vocabulary Remote Sensing Instance SegmentationShiqi Huang, Shuting He, Huaiyuan Qin, Bihan WenICCV 2025 · 被引用 2 次
- Improving Fusion of Region Features and Grid Features via Two-Step Interaction for Image-Text RetrievalDongqing Wu, Huihui Li, Cang Gu, Lei Guo 等ACM MM 2022 · 被引用 10 次
- Revisiting Self-Similarity: Structural Embedding for Image RetrievalSeongwon Lee, Suhyeon Lee, Hongje Seong, Euntai KimCVPR 2023
