DVF: Advancing Robust and Accurate Fine-Grained Image Retrieval with Retrieval Guidelines
Xin Jiang, Hao Tang, Rui Yan, Jinhui Tang, Zechao Li
Abstract
Fine-grained image retrieval (FGIR) is to learn visual representations that distinguish visually similar objects while maintaining generalization. Existing methods propose to generate discriminative features, but rarely consider the particularity of the FGIR task itself. This paper presents a meticulous analysis leading to the proposal of practical guidelines to identify subcategory-specific discrepancies and generate discriminative features to design effective FGIR models. These guidelines include emphasizing the object (G1), highlighting subcategory-specific discrepancies (G2), and employing effective training strategy (G3). Following G1 and G2, we design a novel Dual Visual Filtering mechanism for the plain visual transformer, denoted as DVF, to capture subcategory-specific discrepancies. Specifically, the dual visual filtering mechanism comprises an object-oriented module and a semantic-oriented module. These components serve to magnify objects and identify discriminative regions, respectively. Following G3, we implement a discriminative model training strategy to improve the discriminability and generalization ability of DVF. Extensive analysis and ablation studies confirm the efficacy of our proposed guidelines. Without bells and whistles, the proposed DVF achieves state-of-the-art performance on three widely-used fine-grained datasets in closed-set and open-set settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 12dd0591-d5e8-46ba-ba0c-c1798e5d168eCited by top-tier papers6
- Multi-scale Activation, Selection, and Aggregation: Exploring Diverse Cues for Fine-Grained Bird RecognitionZhicheng Zhang, Hao Tang, Jinhui TangAAAI 2025 · 6 citations
- Tensor-Aggregated LoRA in Federated Fine-TuningZhixuan Li, Binqian Xu, Xiangbo Shu, Jiachao Zhang et al.ICCV 2025 · 2 citations
- FedMGP: Personalized Federated Learning with Multi-Group Text-Visual PromptsWeihao Bo, Yanpeng Sun, Yu Wang, Xinyu Zhang et al.NeurIPS 2025 · 2 citations
- Fine-Grained Image Retrieval via Dual-Vision AdaptationXin Jiang, Meiqi Cao, Hao Tang, Fei Shen et al.AAAI 2026 · 1 citation
- Exploiting Frequency Dynamics for Enhanced Multimodal Event-Based Action RecognitionMeiqi Cao, Xiangbo Shu, Xin Jiang, Rui Yan et al.ICCV 2025 · 1 citation
Builds on15
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
- SIM-Trans: Structure Information Modeling Transformer for Fine-grained Visual CategorizationHongbo Sun, Xiangteng He, Yuxin PengACM MM 2022 · 128 citations
- Hyperbolic Vision Transformers: Combining Improvements in Metric LearningAleksandr Ermolov, Leyla Mirvakhabova, Valentin Khrulkov, Nicu Sebe et al.CVPR 2022 · 97 citations
- Delving into Multimodal Prompting for Fine-Grained Visual ClassificationXin Jiang, Hao Tang, Junyao Gao, Xiaoyu Du et al.AAAI 2024 · 71 citations
Related papers
- Adversarial Reconstruction Feedback for Robust Fine-Grained GeneralizationShijie Wang, Jian Shi, Haojie LiICCV 2025 · 2 citations
- TransFG: A Transformer Architecture for Fine-Grained RecognitionJu He, Jieneng Chen, Shuai Liu, Adam Kortylewski et al.AAAI 2022 · 529 citations
- DiT-Distill: Open-Set Fine-Grained Retrieval via Generative Curriculum KnowledgeXin Jiang, Hao Tang, Meiqi Cao, Junyao Gao et al.CVPR 2026
- Fine-Grained Retrieval Prompt TuningShijie Wang, Jianlong Chang, Zhihui Wang, Haojie Li et al.AAAI 2023 · 27 citations
- Open-Set Fine-Grained Retrieval via Prompting Vision-Language EvaluatorShijie Wang, Jianlong Chang, Haojie Li, Zhihui Wang et al.CVPR 2023
