ARNet: Self-Supervised FG-SBIR with Unified Sample Feature Alignment and Multi-Scale Token Recycling
Jianan Jiang, Hao Tang, Zhilin Jiang, Weiren Yu, Di Wu
摘要
Fine-Grained Sketch-Based Image Retrieval (FG-SBIR) aims to minimize the distance between sketches and corresponding images in the embedding space. However, scalability is hindered by the growing complexity of solutions, mainly due to the abstract nature of fine-grained sketches. In this paper, we propose an effective approach to narrow the gap between the two domains. It mainly facilitates unified mutual information sharing both intra-and inter-samples, rather than treating them as a single feature alignment problem between modalities. Specifically, our approach includes: (i) Employing dual weight-sharing networks to optimize alignment within the sketch and image domain, which also effectively mitigates model learning saturation issues. (ii) Introducing an objective optimization function based on contrastive loss to enhance the model's ability to align features in both intra-and inter-samples. (iii) Presenting a self-supervised Multi-Scale Token Recycling (MSTR) Module featured by recycling discarded patch tokens in multi-scale features, further enhancing representation capability and retrieval performance. Our framework achieves excellent results on CNN-and ViT-based backbones. Extensive experiments demonstrate its superiority over existing methods. We also introduce Cloths-V1, the first professional fashion sketch-image dataset, utilized to validate our method and will be beneficial for other applications. Our code, new datasets, and pre-trained model are available at https://github.com/ExponentiAI/ARNet .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper15
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal 等NeurIPS 2020 · 被引用 5,249 次
- Transformer in TransformerKai Han, An Xiao, Enhua Wu, Jianyuan Guo 等NeurIPS 2021 · 被引用 2,148 次
相关 Paper
- DLA-Net for FG-SBIR: Dynamic Local Aligned Network for Fine-Grained Sketch-Based Image RetrievalJiaqing Xu, Haifeng Sun, Qi Qi, Jingyu Wang 等ACM MM 2021 · 被引用 14 次
- Exploiting Unlabelled Photos for Stronger Fine-Grained SBIRAneeshan Sain, Ayan Kumar Bhunia, Subhadeep Koley, Pinaki Nath Chowdhury 等CVPR 2023
- SketchMind: Understanding Abstract Sketches with MLLMs for Fine-Grained Sketch-Based Image RetrievalChangxing Li, Donglin Zhang, Zhikai Hu, Xiao-Jun Wu 等WWW 2026
- Photo Pre-Training, But for SketchKe Li, Kaiyue Pang, Yi-Zhe SongCVPR 2023
- Asymmetric Mutual Alignment for Unsupervised Zero-Shot Sketch-Based Image RetrievalZhihui Yin, Jiexi Yan, Chenghao Xu, Cheng DengAAAI 2024 · 被引用 6 次
