DLA-Net for FG-SBIR: Dynamic Local Aligned Network for Fine-Grained Sketch-Based Image Retrieval
Jiaqing Xu, Haifeng Sun, Qi Qi, Jingyu Wang, Ce Ge, Lejian Zhang, Jianxin Liao
摘要
Fine-grained sketch-based image retrieval is considered as an ideal alternative to keyword-based image retrieval and image search by image due to the rich and easily accessible characteristics of sketches. Previous works always follow a paradigm that first extracting image global feature with convolution neural network and then optimizing the model with triplet loss. Many efforts on narrowing the domain gap and extracting discriminating features are made by these works. However, they ignored that the global feature is not good at capturing fine-grained details. In this paper, we emphasize the local features are more discriminating than global feature in FG-SBIR and explore an effective way to utilize local features. Specifically, Local Aligned Network (LA-Net) is proposed first, which solves FG-SBIR by directly aligning the mid-level local features. Experiment manifests it can beat all previous baselines and is easy to implement. LA-Net is hoped to be a new strong baseline for FG-SBIR. Next, Dynamic Local Aligned Network (DLA-Net) is proposed to enhance LA-Net. The question of spatial misalignment caused by the abstraction of the sketch is not considered by LA-Net. To solve this question, a dynamic alignment mechanism is introduced into LA-Net. This new mechanism makes the sketch interact with the photo and dynamically decide where to align according to the different photos. The Experiment indicates DLA-Net successfully addresses the question of spatial misalignment. It gains a significant performance boost over LA-Net and outperforms the state-of-the-art in FG-SBIR. To the best of our knowledge, DLA-Net is the first model that beats humans on all datasets---QMUL FG-SBIR, QMUL Handbag, and Sketchy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Neural Image Popularity Assessment with Retrieval-augmented TransformerLiya Ji, Chan Ho Park, Zhefan Rao, Qifeng ChenACM MM 2023 · 被引用 6 次
- Text-to-Image Diffusion Models are Great Sketch-Photo MatchmakersSubhadeep Koley, Ayan Kumar Bhunia, Aneeshan Sain, Pinaki Nath Chowdhury 等CVPR 2024
- You'll Never Walk Alone: A Sketch and Text Duet for Fine-Grained Image RetrievalSubhadeep Koley, Ayan Kumar Bhunia, Aneeshan Sain, Pinaki Nath Chowdhury 等CVPR 2024
它引用的顶会 Paper3
- DeepFacePencil: Creating Face Images from Freehand SketchesYuhang Li, Xuejin Chen, Binxin Yang, Zihan Chen 等ACM MM 2020 · 被引用 56 次
- Coupling Deep Textural and Shape Features for Sketch RecognitionQi Jia, Xin Fan, Meiyu Yu, Yuqing Liu 等ACM MM 2020 · 被引用 12 次
- Solving Mixed-Modal Jigsaw Puzzle for Fine-Grained Sketch-Based Image RetrievalKaiyue Pang, Yongxin Yang, Timothy M. Hospedales, Tao Xiang 等CVPR 2020
相关 Paper
- Multi-Level Region Matching for Fine-Grained Sketch-Based Image RetrievalZhixin Ling, Zhen Xing, Jiangtong Li, Li NiuACM MM 2022 · 被引用 17 次
- ARNet: Self-Supervised FG-SBIR with Unified Sample Feature Alignment and Multi-Scale Token RecyclingJianan Jiang, Hao Tang, Zhilin Jiang, Weiren Yu 等AAAI 2025 · 被引用 4 次
- Exploiting Unlabelled Photos for Stronger Fine-Grained SBIRAneeshan Sain, Ayan Kumar Bhunia, Subhadeep Koley, Pinaki Nath Chowdhury 等CVPR 2023
- Sketch Less for More: On-the-Fly Fine-Grained Sketch-Based Image RetrievalAyan Kumar Bhunia, Yongxin Yang, Timothy M. Hospedales, Tao Xiang 等CVPR 2020
- SketchMind: Understanding Abstract Sketches with MLLMs for Fine-Grained Sketch-Based Image RetrievalChangxing Li, Donglin Zhang, Zhikai Hu, Xiao-Jun Wu 等WWW 2026
