Learning Token-Based Representation for Image Retrieval
Hui Wu, Min Wang, Wengang Zhou, Yang Hu, Houqiang Li
摘要
In image retrieval, deep local features learned in a data-driven manner have been demonstrated effective to improve retrieval performance. To realize efficient retrieval on large image database, some approaches quantize deep local features with a large codebook and match images with aggregated match kernel. However, the complexity of these approaches is nontrivial with large memory footprint, which limits their capability to jointly perform feature learning and aggregation. To generate compact global representations while maintaining regional matching capability, we propose a unified framework to jointly learn local feature representation and aggregation. In our framework, we first extract deep local features using CNNs. Then, we design a tokenizer module to aggregate them into a few visual tokens, each corresponding to a specific visual pattern. This helps to remove background noise, and capture more discriminative regions in the image. Next, a refinement block is introduced to enhance the visual tokens with self-attention and cross-attention. Finally, different visual tokens are concatenated to generate a compact global representation. The whole framework is trained end-to-end with image-level labels. Extensive experiments are conducted to evaluate our approach, which outperforms the state-of-theart methods on the Revisited Oxford and Paris datasets. Our code is available at https://github.com/MCC-WH/Token .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Learning Spatial-context-aware Global Visual Feature Representation for Instance Image RetrievalZhongyan Zhang, Lei Wang, Luping Zhou, Piotr KoniuszICCV 2023 · 被引用 13 次
- Coarse-to-Fine: Learning Compact Discriminative Representation for Single-Stage Image RetrievalYunquan Zhu, Xinkai Gao, Bo Ke, Ruizhi Qiao 等ICCV 2023 · 被引用 8 次
- On Train-Test Class Overlap and Detection for Image RetrievalChull Hwan Song, Jooyoung Yoon, Taebaek Hwang, Shunghyun Choi 等CVPR 2024 · 被引用 3 次
- Asymmetric Feature Fusion for Image RetrievalHui Wu, Min Wang, Wengang Zhou, Zhenbo Lu 等CVPR 2023
- Revisiting Self-Similarity: Structural Embedding for Image RetrievalSeongwon Lee, Suhyeon Lee, Hongje Seong, Euntai KimCVPR 2023
它引用的顶会 Paper3
- Learning With Average Precision: Training Image Retrieval With a Listwise LossJérôme Revaud, Jon Almazán, Rafael S. Rezende, César Roberto de SouzaICCV 2019 · 被引用 424 次
- Learning Deep Local Features with Multiple Dynamic Attentions for Large-Scale Image RetrievalHui Wu, Min Wang, Wengang Zhou, Houqiang LiICCV 2021 · 被引用 26 次
- Google Landmarks Dataset v2 - A Large-Scale Benchmark for Instance-Level Recognition and RetrievalTobias Weyand, André Araújo, Bingyi Cao, Jack SimCVPR 2020
相关 Paper
- Towards Implicit Aggregation: Robust Image Representation for Place Recognition in the Transformer EraFeng Lu, Tong Jin, Canming Ye, Xiangyuan Lan 等NeurIPS 2025 · 被引用 8 次
- DOLG: Single-Stage Image Retrieval with Deep Orthogonal Fusion of Local and Global FeaturesMin Yang, Dongliang He, Miao Fan, Baorong Shi 等ICCV 2021 · 被引用 135 次
- GloTok: Global Perspective Tokenizer for Image Reconstruction and GenerationXuan Zhao, Zhongyu Zhang, Yuge Huang, Yuxi Mi 等AAAI 2026 · 被引用 1 次
- Learning Compact 3D Representations from Feed-Forward Novel View SynthesisHonggyu An, Jaewoo Jung, Mungyeom Kim, Chaehyun Kim 等CVPR 2026
- Improving Fusion of Region Features and Grid Features via Two-Step Interaction for Image-Text RetrievalDongqing Wu, Huihui Li, Cang Gu, Lei Guo 等ACM MM 2022 · 被引用 10 次
